Cognition's SWE-2 coding model nears frontier AI performance at lower cost
New CapabilitiesNew model scores 92.8 on Terminal-Bench 2.1 and comes within a point of Claude Fable 5.1 on FrontierCode, at a claimed 64% lower cost.
Today: SWE-2 unveiled with leading benchmark scoresNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated 2 hours agoCognition released SWE-2, a coding model that scored 92.8 on Terminal-Bench 2.1 and 50.0 on FrontierCode 1.1 Main, within one point of Anthropic's Claude Fable 5.1. Cognition claims SWE-2 costs 64% less than Fable 5.1 for that performance.
The 27.3 score on Terminal-Bench 4.0, released weeks ago, shows long-horizon agentic tasks remain frontier labs' stronghold. If SWE-2 closes that gap, its cost advantage could reshape how enterprises buy coding AI.
Why it matters
If SWE-2 holds up, coding AI costs could fall by up to 75%, reshaping the developer tools market.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
AI company building autonomous coding agents and the SWE model line.
Chinese AI lab that developed the Kimi K3 model, used as the base for SWE-2.
AI lab that produces Claude Fable 5.1, a frontier coding model compared against SWE-2.
AI lab that develops GPT-6 Astra, a frontier coding model that leads on Terminal-Bench 4.0.
Timeline
-
SWE-2 unveiled with leading benchmark scores
Today Product LaunchCognition releases SWE-2, scoring 92.8 on Terminal-Bench 2.1, 50.0 on FrontierCode 1.1 Main, and 73.0 on DeepSWE 1.1, while claiming 64% lower cost than Fable 5.1.
Historical Context
2 moments from history that rhyme with this story — and how they unfolded.
ImageNet saturation (2015–2017)
By 2015, algorithms topped 95% accuracy on ImageNet, leading researchers to create harder benchmarks like COCO and WinoGrande. The pattern repeated in coding: SWE-bench saturated, prompting new tests like Terminal-Bench 4.0.
Models appeared to solve old benchmarks while still failing on harder tasks.
Benchmark developers responded with new tests that separated robust models from overfitted ones.
SWE-2's 92.8 on Terminal-Bench 2.1 contrasts with its 27.3 on Terminal-Bench 4.0, highlighting how newer, harder benchmarks expose gaps that older ones mask.
DeepSeek V3 (December 2024)
DeepSeek released a Mixture-of-Experts model that matched OpenAI's GPT-4 on coding benchmarks at a fraction of the training cost. The model triggered a sell-off in AI chip stocks and forced frontier labs to justify their pricing.
Frontier labs kept performance leads but faced new price pressure; enterprise buyers gained leverage.
The event accelerated optimization of inference costs and popularized reinforcement learning on open base models.
SWE-2 follows the same playbook: post-training an open base model (Kimi K3) to near-frontier performance at lower cost, pressuring incumbents on price.
