AI news ·
Mayo Clinic AI detects pancreatic cancer 16 months early as studies show LLMs outperform physicians in diagnosis
Two AI studies show machines outperforming doctors: Mayo Clinic's model spotted pancreatic cancer 16 months early, while OpenAI's o1 correctly diagnosed complex cases 78% of the time.

Two Studies Show AI Can Detect Cancers Earlier and Diagnose Complex Cases Better Than Doctors
Mayo Clinic researchers demonstrated that their AI model can identify pancreatic cancer 16 months before typical diagnosis, while Harvard and Stanford researchers found that OpenAI's latest language model outperforms physicians in complex diagnostic cases. The findings suggest healthcare systems need to prepare their infrastructure and workflows for AI-assisted clinical decision-making.
Mayo Clinic's AI Detects Pancreatic Cancer Three Times Earlier
Mayo Clinic's Radiomics-based Early Detection Model, or REDMOD, identified pancreatic ductal adenocarcinoma at a median lead time of 475 days-about 16 months-before patients received clinical diagnoses. The model analyzed nearly 2,000 CT scans, including scans originally interpreted as normal from patients later diagnosed with pancreatic cancer.
REDMOD achieved 73% sensitivity compared to 39% for radiologists reviewing the same scans without AI assistance. For cancers detected more than 24 months before diagnosis, the AI identified nearly three times as many early cases that would otherwise go undetected.
The researchers called the result "a necessary step towards shifting the paradigm from late-stage symptomatic diagnosis to proactive pre-clinical interception." Pancreatic cancer remains one of the deadliest cancers partly because it typically goes undetected until advanced stages when treatment options are limited.
Advanced Language Model Outperforms Physicians in Emergency Settings
Researchers at Harvard Medical School, Beth Israel Deaconess Medical Center, and Stanford University tested OpenAI's o1-series model on challenging clinical cases similar to those presented in the New England Journal of Medicine's diagnostic conference series, which has served as a benchmark for diagnostic reasoning since the 1950s.
The model found the correct diagnosis in its differential list 78.3% of the time. In 52% of cases, the correct diagnosis was its first suggestion. When researchers included the model's "potentially helpful" or "very close" diagnoses, accuracy surged to 97.9%.
When tested against GPT-4, the o1-preview model achieved 88.6% accuracy compared to GPT-4's 72.9%. The o1-preview outperformed GPT-4 in 24.3% of cases, while GPT-4 exceeded it in only 7.1% of cases.
The researchers compared the model's performance to baseline physician performance using real, unstructured clinical data from an emergency department at a major Boston academic hospital. The model showed the widest performance gap over physicians at initial emergency room triage, where clinicians have the least information available.
Healthcare Systems Must Prepare Infrastructure and Workflows
The researchers said their findings indicate healthcare systems need to invest in computing infrastructure and design workflows that support "clinician-AI interaction" to safely integrate these tools into patient care.
Effective implementation requires robust monitoring frameworks that track not just diagnostic accuracy but also safety, efficiency, and cost. The researchers acknowledged their study had limitations but said rapid improvements in language models hold substantial implications for clinical medicine and indicate an urgent need for real-world evaluation.
While applying AI to clinical decision support carries risks, the researchers argued that greater use of these tools might reduce the human and financial costs of diagnostic error, delay, and lack of access to specialist expertise.
Learn more about AI for Healthcare and Generative AI and LLM applications in clinical settings.