DeepSeek releases V4.1-Flash, a cheaper model for long-context AI agents
New CapabilitiesNew architecture cuts per-token memory to a quarter of the prior generation; V4-Pro is set to retire
Today: DeepSeek announces V4.1-Flash architecture detailsNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated 49 minutes agoDeepSeek released V4.1-Flash on September 10, a model that reads a million tokens of context while using a fraction of the memory its predecessors needed. The key-value cache, the memory a model keeps to recall what it has already read, drops to 890 bytes per token, one quarter of the prior Flash generation.
The price is the point. Off-peak, cached input runs $0.003 per million tokens, half the peak rate. DeepSeek says internal tests show the new model beats its larger V4-Pro line on cost, speed, and quality, and it plans to route V4-Pro requests to V4.1-Flash starting September 14.
Why it matters
Off-peak cached input costs $0.003 per million tokens, making always-on AI agents cheap enough to run at scale.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
Timeline
-
DeepSeek routes V4-Pro traffic to V4.1-Flash
Upcoming Migrationdeepseek-v4-pro requests begin serving V4.1-Flash at Flash rates. Expected to continue until V4.1-Pro ships.
-
DeepSeek announces V4.1-Flash architecture details
Today AnnouncementDetails the causal encoder-decoder design, 890-byte-per-token FP4 cache, and September 14 V4-Pro retirement.
-
V4.1-Flash goes live on DeepSeek API
ReleaseModel launches as deepseek-flash with new peak and off-peak pricing. V4-Flash requests now route to it.
Historical Context
3 moments from history that rhyme with this story — and how they unfolded.
DeepSeek V2's Multi-head Latent Attention (May 2024)
DeepSeek-V2 introduced Multi-head Latent Attention, which compressed the key-value cache into a low-rank latent space. The technique cut cache memory sharply and let the model serve longer contexts on the same hardware.
V2 cut inference costs and boosted throughput relative to comparable-size models of its day.
MLA was widely copied across the industry and became the template for cache-efficient attention.
V4.1-Flash's 890-byte-per-token cache is the direct descendant of that design, pushing the same obsession from training efficiency into serving efficiency.
Google Gemini 1.5's million-token context (February 2024)
Google announced Gemini 1.5 Pro with a 1-million-token context window, the first mainstream model to handle that much input at once.
The announcement reset expectations for what long-context AI could process, from entire books to full codebases.
Million-token context became a benchmark that major labs now race to match; the differentiator shifted to how cheaply it can be served.
DeepSeek's contribution is not the 1M window itself, but making that window affordable enough for continuous agent workloads.
DeepSeek R1 market shock (January 2025)
DeepSeek released R1, a reasoning model trained for a reported fraction of what rivals spent, matching frontier performance on several benchmarks. On January 27, 2025, Nvidia lost about $589 billion in market value as investors questioned whether frontier AI required as many chips as projected.
R1 topped app download charts worldwide and forced an industry-wide reassessment of training costs.
Reset the assumption that frontier AI required billion-dollar training runs, and made efficiency a first-class goal for every lab.
V4.1-Flash runs the same playbook, but targets inference memory cost rather than training cost.
