Microsoft's experimental AI diagnostic system, paired with OpenAI's o3 model, correctly solved up to 85.5% of 304 exceptionally difficult medical cases adapted from the New England Journal of Medicine - while 21 experienced physicians averaged just 19.9% under the benchmark conditions. The June 2025 preprint result made headlines, but a major revision published in November 2025 changed the numbers: maximum accuracy of 84.5% for the AI and an average of 36.1% for physicians. The result remains an impressive demonstration of structured AI reasoning, not proof the system can replace doctors.
The work is research reporting, not medical advice, and the system is not a substitute for clinical care. It was tested on a controlled, text-based benchmark, not on real patients. The original manuscript has not completed journal peer review, and the revised figures show how much a preprint can change.
What the benchmark actually tested
Researchers built a benchmark called SDBench from 304 consecutive clinicopathological conference cases published between 2017 and 2025. These are not routine complaints from a waiting room. The New England Journal of Medicine's Case Records are educational investigations that revolve around rare conditions, ambiguous evidence, or surprising diagnoses. That selection makes the exercise useful for testing diagnostic breadth, but it changes how accuracy should be interpreted. Performance on a collection of difficult, unusual cases does not reveal how a system would behave when common conditions dominate.
SDBench converted each published case into a sequence. The diagnostic system started with limited information, requested tests, revised its differential diagnosis, and committed to an answer. A separate gatekeeper model held the complete case and disclosed information in response to those requests. That made the task more demanding than selecting an answer from a list, while remaining very different from bedside medicine. The most recent 56 cases formed a hidden test set, and physicians were evaluated on cases from that set.
How one model acted like a panel
The Microsoft AI Diagnostic Orchestrator, MAI-DxO, did not simply ask a chatbot for a diagnosis. It directed one underlying language model to take on five jobs: a hypothesis agent proposed possible diagnoses, a challenger searched for weaknesses, a checklist agent looked for missing evidence, a test-selection agent chose the next investigation, and a stewardship agent weighed whether more testing was justified. Those roles created a structured loop of proposing, criticizing, investigating, and revising.
Accuracy came with substantial simulated cost. The authors said the maximum-accuracy configuration used about $7,184 in model and test costs across the benchmark. A more frugal configuration reached 79.9% for roughly $2,396. Performance depended in part on how much deliberation and testing the system was allowed.
Why 85.5 versus 20 is not a bedside contest
The physician group included 17 primary-care professionals and four hospital workers from the United States and United Kingdom, with a median 12 years of experience. Each doctor reasoned alone without search engines or reference databases, while MAI-DxO simulated a multi-role panel. The AI handled all 304 cases; the human comparison came from the held-out subset. The paper also used different interaction formats, though the authors reported that switching the AI to the physicians' single-turn format did not reduce its accuracy on the test set.
Both sides encountered curated text. There was no physical examination, live patient conversation, imaging review, or fragmented electronic record. A correct name for a rare diagnosis is valuable, but clinical practice also requires deciding what is urgent, explaining uncertainty, and avoiding harm.
The November 2025 revision broadened the evaluation by adding 32 emergency-department cases from MIMIC-IV-ED. Its maximum accuracy was 84.5% on the journal cases and 80.9% on the emergency set. Physician averages were 36.1% and 46.8%, respectively. The revised protocol, expanded dataset, and updated model can change outcomes - the attached date matters. Preprints are useful because results can be examined immediately, but they can also shift before publication.
What the result still demonstrates
Even with those caveats, the experiment offers a technical lesson. Breaking diagnostic reasoning into explicit roles can outperform a one-shot answer from the same class of model. Orchestration, criticism, and disciplined data collection may matter almost as much as raw model capability. It also shows why a strong standalone score is only the beginning. A 2024 randomized clinical trial in JAMA Network Open found that giving 50 physicians access to GPT-4 did not significantly improve their diagnostic reasoning over conventional resources, even though the model alone scored better in an exploratory comparison.
There are more clinically grounded ways to evaluate medical AI. A large screening study, previously covered by ScienceBlog, measured cancer detection and radiologist workload inside a real care program. Such studies answer different questions from a benchmark of rare diagnostic puzzles.
Why this matters for science and research professionals
Professionals working in medical research or AI evaluation should treat this result as a system design example, not a clinical ready-for-primetime claim. The gap between scripted performance and bedside usefulness is large, and the comparison between physicians and AI on deliberate benchmarks does not reflect how either performs in the world. Those pursuing AI in medical research should look for prospective trials with representative patients, real clinical teams, and tracking of outcomes like missed diagnoses, unnecessary tests, timeliness, fairness, and harm - not just accuracy scores on curated cases. The productive question is whether carefully evaluated tools can help clinical teams reach safer decisions, while leaving responsibility and patient care firmly anchored in medicine.
Your membership also unlocks: