Pull to refresh
Logo
DeepSeek releases V4.1-Flash, a cheaper model for long-context AI agents

DeepSeek releases V4.1-Flash, a cheaper model for long-context AI agents

New Capabilities

New architecture cuts per-token memory to a quarter of the prior generation; V4-Pro is set to retire

Today: DeepSeek announces V4.1-Flash architecture details

Overview

Updated 49 minutes ago

DeepSeek released V4.1-Flash on September 10, a model that reads a million tokens of context while using a fraction of the memory its predecessors needed. The key-value cache, the memory a model keeps to recall what it has already read, drops to 890 bytes per token, one quarter of the prior Flash generation.

The price is the point. Off-peak, cached input runs $0.003 per million tokens, half the peak rate. DeepSeek says internal tests show the new model beats its larger V4-Pro line on cost, speed, and quality, and it plans to route V4-Pro requests to V4.1-Flash starting September 14.

Why it matters

Off-peak cached input costs $0.003 per million tokens, making always-on AI agents cheap enough to run at scale.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

890 bytes/token
Key-value cache per token
Global cache footprint, one quarter of V4-Flash's and one eighth of its persistent SSD cache.
1 million tokens
Context window
Native input context, with maximum output of 384,000 tokens.
$0.003
Off-peak cache-hit input price per million tokens
Half the peak rate, applying outside weekday 01:00-04:00 and 06:00-10:00 UTC windows.
45 trillion
Training tokens
Multimodal corpus the model was trained on from scratch.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

Play

Exploring all sides of a story is often best achieved with Play.

Most of these play right now — no account needed. Sign up to save scores, keep a streak, and unlock Debate and Predict. Log in Sign Up
Predict 3 ways this could play out. Back the one you believe — contrarian picks score more when a scenario has a resolution date. Log in to play

People Involved

Organizations Involved

Timeline

3 events Latest: Today
  1. DeepSeek routes V4-Pro traffic to V4.1-Flash

    Upcoming Migration

    deepseek-v4-pro requests begin serving V4.1-Flash at Flash rates. Expected to continue until V4.1-Pro ships.

  2. DeepSeek announces V4.1-Flash architecture details

    Today Announcement

    Details the causal encoder-decoder design, 890-byte-per-token FP4 cache, and September 14 V4-Pro retirement.

  3. V4.1-Flash goes live on DeepSeek API

    Release

    Model launches as deepseek-flash with new peak and off-peak pricing. V4-Flash requests now route to it.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

May 2024

DeepSeek V2's Multi-head Latent Attention (May 2024)

DeepSeek-V2 introduced Multi-head Latent Attention, which compressed the key-value cache into a low-rank latent space. The technique cut cache memory sharply and let the model serve longer contexts on the same hardware.

Then

V2 cut inference costs and boosted throughput relative to comparable-size models of its day.

Now

MLA was widely copied across the industry and became the template for cache-efficient attention.

Why this matters now

V4.1-Flash's 890-byte-per-token cache is the direct descendant of that design, pushing the same obsession from training efficiency into serving efficiency.

February 2024

Google Gemini 1.5's million-token context (February 2024)

Google announced Gemini 1.5 Pro with a 1-million-token context window, the first mainstream model to handle that much input at once.

Then

The announcement reset expectations for what long-context AI could process, from entire books to full codebases.

Now

Million-token context became a benchmark that major labs now race to match; the differentiator shifted to how cheaply it can be served.

Why this matters now

DeepSeek's contribution is not the 1M window itself, but making that window affordable enough for continuous agent workloads.

January 2025

DeepSeek R1 market shock (January 2025)

DeepSeek released R1, a reasoning model trained for a reported fraction of what rivals spent, matching frontier performance on several benchmarks. On January 27, 2025, Nvidia lost about $589 billion in market value as investors questioned whether frontier AI required as many chips as projected.

Then

R1 topped app download charts worldwide and forced an industry-wide reassessment of training costs.

Now

Reset the assumption that frontier AI required billion-dollar training runs, and made efficiency a first-class goal for every lab.

Why this matters now

V4.1-Flash runs the same playbook, but targets inference memory cost rather than training cost.

Sources

(7)