AI clinical notes hide dangerous errors that automated checks miss

About 1 in 20 AI-generated clinical notes contains an error serious enough to harm a patient, and the systems are already used in a third of U.S. practices. The most dangerous failures are subtle omissions, like missing a jaw-ache symptom that signals a blinding condition.

Categorized in: AI News Healthcare
Published on: Aug 23, 2026
AI clinical notes hide dangerous errors that automated checks miss

AI-generated clinical notes are now used in roughly a third of U.S. medical practices, but the systems producing them carry a hidden risk: errors that look correct on the surface can be the most dangerous. Sebastian Fox, a medical doctor and CEO of Componso, an AI evaluation company, highlighted this problem in a recent presentation, pointing to real-world data showing that about 1 in 20 AI-generated notes contains an error serious enough to potentially harm a patient.

When a correct note is dangerously wrong

Fox opened with a case that illustrates the subtlety of the problem. An AI-generated note for a patient with a headache recommended paracetamol for a tension-type headache. The note read as routine and technically accurate. But it omitted a detail the patient had mentioned: jaw ache when chewing. For a patient over 50, that combination points to giant cell arteritis, a condition that can cause blindness within days if untreated.

"The most dangerous AI errors are often the ones that appear perfectly fine on the surface," Fox said. The note treated a potential emergency as a common ailment because the AI failed to weigh the importance of one specific detail.

He contrasted that with a more obvious failure: an AI incorrectly diagnosed a young man with diabetes and prescribed medications that don't exist, leading to an invitation for diabetic eye screening for a condition he didn't have. Glaring errors like that tend to get caught. The quieter ones, Fox argued, slip through and can cause greater harm.

The scale of the problem

Fox cited a real-world study showing that nearly 1 in 5 AI-generated clinical notes had an important omission, and more than 1 in 10 contained a hallucination. With ambient scribes already deployed in about a third of U.S. practices, these numbers represent a significant volume of patient encounters.

The lack of comprehensive adverse event reporting for most AI systems means these errors often go undetected. Fox described this as a "flying blind" scenario where the impact on patient care is unknown but potentially severe. He noted the issue extends beyond healthcare to any high-stakes application of AI.

Why current evaluation falls short

Fox explained that large language models now produce fewer obvious factual errors. The failures have shifted to subtle misinterpretations and omissions. Automated checks catch some of the obvious problems, but many serious ones slip through because they hinge on context.

The core issue, he said, is that AI struggles with "what matters" in a given situation. This discernment is tacit, contextual, and constantly evolving, which makes it hard to codify into a fixed rubric. AI can process information, but it lacks the human judgment to distinguish critical details from routine ones. This is a fundamental challenge for Generative AI and LLM systems in clinical settings.

A continuous evaluation loop

Fox proposed a solution built on a continuous feedback loop rather than static evaluation methods. The system has three components: discover failure modes from real-world outputs, capture expert judgment on those failures, and calibrate context-specific standards by referencing past judgments, corrections, and guidelines.

The approach keeps "taste" as examples themselves, allowing the system to learn and adapt over time. Fox said the easiest way to begin is to have experts provide free-form comments on real outputs, which then serve as the raw material for iterative improvement.

Why this matters for healthcare professionals

For clinicians working with AI-generated notes, the practical takeaway is to treat these tools as drafts requiring review, not finished products. The study data Fox cited suggests that reviewing for omissions - not just factual errors - is where the risk concentrates. A note can be technically correct and still miss the detail that changes the diagnosis. That distinction matters for anyone whose name goes on a clinical note.

The evaluation loop Fox describes also points to a broader shift for AI for Healthcare: the systems that work best are those that improve through continuous expert feedback, not one-time validation. Clinicians who flag subtle errors in real outputs are not just fixing individual notes - they're contributing to the calibration that makes these tools safer over time.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)