Pull to refresh
Logo
Ex-DeepMind team launches Faraday AI claiming it beats frontier models at paper replication

Ex-DeepMind team launches Faraday AI claiming it beats frontier models at paper replication

New Capabilities

A 27B-parameter agent from London startup Inherent reportedly out-replicated Claude Opus 4.8 and GPT-5.5 — results are self-reported and unverified

Today: Faraday launch gains wide coverage and scrutiny

Overview

Updated 42 minutes ago

A London startup founded by four former Google DeepMind researchers says its Faraday AI agent reproduced published research findings more accurately than Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. Faraday runs on Qwen 3.6, a model with just 27 billion parameters — far smaller than the frontier systems it claims to beat.

The results come from Inherent's own paper and its own benchmark, so no outside group has verified them. But the architecture is the novel part: a small 'scientist' model that plans experiments and directs a much larger coding agent to execute them. If it works, it suggests agent design matters more than raw model size for scientific work.

Why it matters

If Faraday's architecture holds up, scientific AI won't need frontier-scale models — small directors can run big tool agents for research.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

27B
Faraday model parameters
Core 'scientist' model, based on Qwen 3.6 — roughly 200 times smaller than the coding agent it directs.
310
Replication tasks in Replica benchmark
Figure-reproduction tasks drawn from 100 research papers across machine learning and AI-for-science domains.
$50M
Seed funding raised by Inherent
Co-led by Index Ventures and Radical Ventures; NVentures (NVIDIA's venture arm) also participated.
73%
In-distribution win rate vs. Claude Opus 4.8
Share of machine learning replication tasks where Faraday produced more faithful results, per the company's paper.
60%
Held-out win rate on AI-for-science tasks
Faraday's reported advantage on physics, biology, and materials science papers it hadn't seen during training.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

May 2026 September 2026

3 events Latest: Today
  1. Faraday launch gains wide coverage and scrutiny

    Today Media coverage

    Tech and science outlets report the claims; commentators note results are self-reported and unverified.

  2. Faraday paper and Replica benchmark released

    Publication

    Preprint reports Faraday outperforms Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.

  3. Inherent exits stealth with $50M seed round

    Funding

    Index Ventures and Radical Ventures co-led the round; NVentures, Ex/Ante, Metaplanet, and others joined.

Scenarios

1

Independent benchmark confirms Faraday's replication advantage

Possible Resolves by Jun 1, 2027

Discussed by: ML researchers who study agent evaluation; the wider scientific AI community watching for external validation

An independent research group — a university lab or another AI lab — obtains access to Faraday or reimplements its training approach on Replica's public tasks. They publish results confirming that a small 'scientist' model directing a large coding agent outperforms Claude Opus 4.8 and GPT-5.5 on replication benchmarks. This would move Faraday from self-reported claim to established result and accelerate adoption of the CAT architecture.

2

Faraday claims challenged; outside groups fail to reproduce the advantage

Possible Resolves by Jun 1, 2027

Discussed by: Laura Martel's analysis of the preprint; industry observers who note the benchmark was built and scored by Inherent itself

Independent groups run the Replica tasks with Claude Opus 4.8 and GPT-5.5 under their own protocols and find the performance gap is smaller than reported, or that Faraday's advantage disappears under different evaluation conditions. The paper's rubric-based judge — built by the same team — faces particular scrutiny. The story recedes as another self-reported lab result without external validation.

3

Journals adopt AI-based pre-publication replication checks

Possible Resolves by Jan 1, 2028

Discussed by: Inherent's own framing of Faraday's use cases; editors at journals facing reproducibility pressure

Whatever happens to Faraday specifically, the broader insight — that AI agents can automatically check whether a paper's figures and results are reproducible — gains traction. One or more major scientific journals launch a formal pilot using AI agents to verify submitted papers' reproducibility before publication. Inherent positions Faraday as the reference system for this workflow.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

August 2015

Open Science Collaboration replication project (2015)

A consortium of 270 psychologists tried to replicate 100 studies from top psychology journals. Only 36% of replications produced statistically significant results, versus 97% of the original studies.

Then

Triggered the 'replication crisis' — intense methodological scrutiny across psychology, economics, and medicine.

Now

Led to registered reports, open-data mandates, and a permanent push for reproducibility checks in scientific publishing.

Why this matters now

Faraday's stated purpose is to automate this kind of replication checking. The replication crisis is the problem the agent claims to address at scale.

November-December 2020

AlphaFold 2 (2020)

DeepMind's AlphaFold 2 predicted protein structures with atomic-level accuracy at CASP14, an open competition where international teams submit blind predictions. The system was validated against targets nobody had seen, not by DeepMind's own benchmark.

Then

Shifted expectations for AI in scientific discovery almost overnight; the AlphaFold Protein Structure Database now serves millions of researchers.

Now

Set a template: AI-for-science claims earn credibility through external, blind evaluation — not self-reported results.

Why this matters now

Faraday comes from the same DeepMind talent pool and claims scientific capability, but its results come from the company's own benchmark and rubric-based judge. The AlphaFold comparison frames what independent validation would look like.

August 2024

Sakana AI's AI Scientist (2024)

Sakana AI and University of Oxford researchers built an end-to-end 'AI Scientist' agent that generated research ideas, wrote code, ran experiments, and produced papers. The output was widely judged incremental, and evaluation was self-referential.

Then

Demonstrated that autonomous research agents are technically feasible — but also that they can generate volume without depth.

Now

Established a baseline expectation for what AI research agents can do, and what critics will say when results are self-reported.

Why this matters now

Faraday's creators positioned their work as a step beyond Sakana's — starting from the harder task of replication rather than original paper generation, and using a novel small-directs-large architecture instead of a single monolithic agent.

Sources

(8)