Pull to refresh
Logo
OpenEvidence launches Darwin, first medical AI to ace US licensing benchmark

OpenEvidence launches Darwin, first medical AI to ace US licensing benchmark

New Capabilities

Four-model family spans five-second answers to deep research reports

Yesterday: OpenEvidence launches four medical AI models

Overview

Updated Yesterday

OpenEvidence released four medical AI models on September 5. The flagship, Darwin, is the first to score 100% on MedQA, a benchmark drawn from United States Medical Licensing Examination (USMLE)-style questions. Three production models answer clinicians in five seconds to five minutes, matching response time to question depth.

The company says more American physicians use its platform than all other AI platforms combined, and it is free to verified US clinicians. Darwin is in research preview, restricted to institutional partners and academic researchers. OpenEvidence says its reasoning will flow into the production models as safety safeguards validate.

Why it matters

The medical AI more US physicians use than any rival now has a perfect exam-scorer, with its reasoning bound for the exam room.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

100%
Darwin score on MedQA
First perfect score on the independent medical AI benchmark.
660/660
MedQA questions answered correctly
Darwin answered every question on the USMLE-style exam set correctly.
Under 5 seconds
Osler response time
Fastest production model, built for point-of-care questions.
~30 seconds
Sackett response time
Mid-tier model with additional rounds of searching.
~5 minutes
Snow response time
Deep-investigation model for complex cases and unsettled evidence.
4
Models released
Darwin in research preview; Osler, Sackett, Snow open to all clinicians.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

Play

Exploring all sides of a story is often best achieved with Play.

Most of these play right now — no account needed. Sign up to save scores, keep a streak, and unlock Debate and Predict. Log in Sign Up
Predict 3 ways this could play out. Back the one you believe — contrarian picks score more when a scenario has a resolution date. Log in to play

People Involved

Organizations Involved

Timeline

June 2026 September 2026

3 events Latest: Yesterday
  1. OpenEvidence launches four medical AI models

    Latest Product Launch

    Darwin tops MedQA with the first perfect score; Osler, Sackett, and Snow open free to all clinicians on web and mobile.

  2. STAT previews the model family

    Media Coverage

    STAT Health Tech reports on the upcoming release and generative AI medical devices reaching the market quickly.

  3. Nature Medicine study questions specialized clinical tools

    Research

    Independent study found general-purpose frontier models outperformed an earlier OpenEvidence tool and UpToDate Expert AI on benchmarks and real clinical queries. It does not evaluate Darwin.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

2013-2022

IBM Watson for Oncology (2013-2022)

IBM's medical AI, built after Watson won Jeopardy, was sold to hospitals as a cancer treatment advisor. MD Anderson cancelled a $62 million project in 2017 after internal documents showed the system produced unsafe treatment recommendations.

Then

Maimonides Medical Center stopped using it; IBM sold off Watson Health in 2022.

Now

Became the cautionary tale for medical AI that dazzles in demos but fails in real clinical settings.

Why this matters now

Darwin faces the same question IBM Watson could not answer: whether benchmark mastery survives contact with real patients and real clinician judgment.

July 2021

DeepMind's AlphaFold (2021)

DeepMind's specialized model predicted protein structures from amino acid sequences, solving a 50-year biology problem that general approaches could not crack, validated independently at CASP14.

Then

The prediction database became a standard research tool used by hundreds of thousands of scientists.

Now

Proved a single-purpose model can beat general ones at a hard scientific task.

Why this matters now

Supports the argument that a dedicated medical reasoning model like Darwin can outperform general frontier models in clinical domains.

March 2023

GPT-4 tops the USMLE (2023)

OpenAI's GPT-4 scored in roughly the 90th percentile on USMLE-style questions in a widely cited study, showing general-purpose models could approach medical exam competence without medical training.

Then

Spurred a wave of medical AI startups and regulatory review of AI clinical tools.

Now

Set the benchmark arms race Darwin now resets with a perfect score.

Why this matters now

Darwin's 100% on MedQA is the direct next step in a benchmark progression general models began in 2023.

Sources

(8)