OpenEvidence has released a group of four medical AI models, including one called Darwin that scored 100% on the MedQA benchmark - a test built from U.S. Medical Licensing Examination-style questions. The launch intensifies competition among developers racing to build tools that can handle the complexity of clinical decision-making, even as experts caution that benchmark performance alone does not capture how these systems work alongside physicians in practice.
Darwin is designed for clinical reasoning on the most difficult cases and research questions. It is currently available only as a research preview to institutional partners, research collaborators, and academic researchers for specific applications. OpenEvidence tested the model on four independent medical AI benchmarks.
On the MedQA dataset, Darwin answered all 660 questions correctly. Claude Fable 5 scored 99.7%, GPT-5.6 Sol reached 99.1%, and Gemini 3.7 Flash hit 99.2%. Darwin also outperformed those three models on the MedXpertQA, HealthBench Professional, and NOHARM benchmarks.
A gap between benchmarks and bedside
OpenEvidence was direct about the limits of these results. The benchmarks evaluate models working alone, with no human in the loop. The company's blog post said that setup "does not match the reality of how clinical decision support tools are used today."
"In practice these tools are an aid to clinical judgment rather than a substitute for it, and an answer is 'good' if it helps a physician make a better decision," the post reads. OpenEvidence cited Feng et al. (2026) as setting the standard the company believes clinical tools should ultimately be judged against.
That framing echoes what clinical AI experts have told Healthtech Analytics: even when AI shows strong potential on clinical reasoning tasks, the goal should be optimal human-AI collaboration, not full automation.
Three additional models tuned for speed and depth
Alongside Darwin, OpenEvidence released Osler, Sackett, and Snow on its platform. Osler is an upgrade to the model currently powering the company's clinical intelligence platform and will become the new default. It returns answers to clinical questions at the point of care in under five seconds.
Sackett runs more rounds of searching than Osler and is more likely to ask for clarification before responding. Answers take roughly 30 seconds to generate. Snow conducts multiple search rounds and produces a full report, with a turnaround time of about five minutes. The company designed Snow for complex cases involving differentials, competing comorbidities, or questions where the evidence base is not well established.
The four-model release follows a busy stretch for OpenEvidence. In the first half of the year, the company introduced a quality grading feature for its AI-generated answers, added a tool that predicts whether a patient has structural heart disease, and launched a coding suggestion capability. These moves reflect the broader push across AI for Healthcare to move from experimental models to tools embedded in clinical workflows.
Why this matters for healthcare professionals
Benchmark scores grab headlines, but the real test for clinical AI is whether it helps a physician make a better decision - not whether it can ace an exam on its own. The speed tiers OpenEvidence built into Osler, Sackett, and Snow suggest the company is thinking about how these tools fit into actual clinical rhythms, from a five-second lookup at the point of care to a five-minute deep dive on a complicated case. For healthcare teams evaluating AI tools, the question is less about which model scores highest and more about which one integrates cleanly into the moment a decision needs to be made.
Your membership also unlocks: