Pull to refresh
Logo
Cognition's SWE-2 coding model nears frontier AI performance at lower cost

Cognition's SWE-2 coding model nears frontier AI performance at lower cost

New Capabilities

New model scores 92.8 on Terminal-Bench 2.1 and comes within a point of Claude Fable 5.1 on FrontierCode, at a claimed 64% lower cost.

Today: SWE-2 unveiled with leading benchmark scores

Overview

Updated 2 hours ago

Cognition released SWE-2, a coding model that scored 92.8 on Terminal-Bench 2.1 and 50.0 on FrontierCode 1.1 Main, within one point of Anthropic's Claude Fable 5.1. Cognition claims SWE-2 costs 64% less than Fable 5.1 for that performance.

The 27.3 score on Terminal-Bench 4.0, released weeks ago, shows long-horizon agentic tasks remain frontier labs' stronghold. If SWE-2 closes that gap, its cost advantage could reshape how enterprises buy coding AI.

Why it matters

If SWE-2 holds up, coding AI costs could fall by up to 75%, reshaping the developer tools market.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

92.8
Terminal-Bench 2.1 score
SWE-2's score on the terminal agent benchmark, the highest in Cognition's published table.
50.0
FrontierCode 1.1 Main score
One point below Claude Fable 5.1 (50.9) and 3.3 below GPT-6 Astra (53.3).
64%
Cost reduction claim vs Fable 5.1
Cognition says SWE-2 is 64% cheaper than Fable 5.1 for similar benchmark performance.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

Play

Exploring all sides of a story is often best achieved with Play.

Most of these play right now — no account needed. Sign up to save scores, keep a streak, and unlock Debate and Predict. Log in Sign Up
Predict 3 ways this could play out. Back the one you believe — contrarian picks score more when a scenario has a resolution date. Log in to play

People Involved

Organizations Involved

Timeline

1 event Latest: Today
  1. SWE-2 unveiled with leading benchmark scores

    Today Product Launch

    Cognition releases SWE-2, scoring 92.8 on Terminal-Bench 2.1, 50.0 on FrontierCode 1.1 Main, and 73.0 on DeepSWE 1.1, while claiming 64% lower cost than Fable 5.1.

Historical Context

2 moments from history that rhyme with this story — and how they unfolded.

2015–2017

ImageNet saturation (2015–2017)

By 2015, algorithms topped 95% accuracy on ImageNet, leading researchers to create harder benchmarks like COCO and WinoGrande. The pattern repeated in coding: SWE-bench saturated, prompting new tests like Terminal-Bench 4.0.

Then

Models appeared to solve old benchmarks while still failing on harder tasks.

Now

Benchmark developers responded with new tests that separated robust models from overfitted ones.

Why this matters now

SWE-2's 92.8 on Terminal-Bench 2.1 contrasts with its 27.3 on Terminal-Bench 4.0, highlighting how newer, harder benchmarks expose gaps that older ones mask.

December 2024

DeepSeek V3 (December 2024)

DeepSeek released a Mixture-of-Experts model that matched OpenAI's GPT-4 on coding benchmarks at a fraction of the training cost. The model triggered a sell-off in AI chip stocks and forced frontier labs to justify their pricing.

Then

Frontier labs kept performance leads but faced new price pressure; enterprise buyers gained leverage.

Now

The event accelerated optimization of inference costs and popularized reinforcement learning on open base models.

Why this matters now

SWE-2 follows the same playbook: post-training an open base model (Kimi K3) to near-frontier performance at lower cost, pressuring incumbents on price.

Sources

(10)