OpenEvidence launches Darwin, first medical AI to ace US licensing benchmark
New CapabilitiesFour-model family spans five-second answers to deep research reports
Yesterday: OpenEvidence launches four medical AI modelsNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated YesterdayOpenEvidence released four medical AI models on September 5. The flagship, Darwin, is the first to score 100% on MedQA, a benchmark drawn from United States Medical Licensing Examination (USMLE)-style questions. Three production models answer clinicians in five seconds to five minutes, matching response time to question depth.
The company says more American physicians use its platform than all other AI platforms combined, and it is free to verified US clinicians. Darwin is in research preview, restricted to institutional partners and academic researchers. OpenEvidence says its reasoning will flow into the production models as safety safeguards validate.
Why it matters
The medical AI more US physicians use than any rival now has a perfect exam-scorer, with its reasoning bound for the exam room.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
AI-powered medical knowledge platform serving US clinicians.
US nonprofit supporting people with rare diseases.
Timeline
June 2026 September 2026
-
OpenEvidence launches four medical AI models
Latest Product LaunchDarwin tops MedQA with the first perfect score; Osler, Sackett, and Snow open free to all clinicians on web and mobile.
-
STAT previews the model family
Media CoverageSTAT Health Tech reports on the upcoming release and generative AI medical devices reaching the market quickly.
-
Nature Medicine study questions specialized clinical tools
ResearchIndependent study found general-purpose frontier models outperformed an earlier OpenEvidence tool and UpToDate Expert AI on benchmarks and real clinical queries. It does not evaluate Darwin.
Historical Context
3 moments from history that rhyme with this story — and how they unfolded.
IBM Watson for Oncology (2013-2022)
IBM's medical AI, built after Watson won Jeopardy, was sold to hospitals as a cancer treatment advisor. MD Anderson cancelled a $62 million project in 2017 after internal documents showed the system produced unsafe treatment recommendations.
Maimonides Medical Center stopped using it; IBM sold off Watson Health in 2022.
Became the cautionary tale for medical AI that dazzles in demos but fails in real clinical settings.
Darwin faces the same question IBM Watson could not answer: whether benchmark mastery survives contact with real patients and real clinician judgment.
DeepMind's AlphaFold (2021)
DeepMind's specialized model predicted protein structures from amino acid sequences, solving a 50-year biology problem that general approaches could not crack, validated independently at CASP14.
The prediction database became a standard research tool used by hundreds of thousands of scientists.
Proved a single-purpose model can beat general ones at a hard scientific task.
Supports the argument that a dedicated medical reasoning model like Darwin can outperform general frontier models in clinical domains.
GPT-4 tops the USMLE (2023)
OpenAI's GPT-4 scored in roughly the 90th percentile on USMLE-style questions in a widely cited study, showing general-purpose models could approach medical exam competence without medical training.
Spurred a wave of medical AI startups and regulatory review of AI clinical tools.
Set the benchmark arms race Darwin now resets with a perfect score.
Darwin's 100% on MedQA is the direct next step in a benchmark progression general models began in 2023.
