A patient arrived at a hospital late one night with a high fever and rapid heart rate. The electronic health record's embedded sepsis prediction model - the same software installed at hundreds of U.S. hospitals - missed the warning signs. Hours later, staff realized the patient was in severe decline, and they had no time to react. This failure is not hypothetical. In 2021, researchers from the University of Michigan published an external study in JAMA Internal Medicine showing that the Epic Systems sepsis model identified only about one-third of patients who would ultimately develop sepsis. The finding exposed a gap between the performance numbers that win FDA clearance and how AI devices actually work in clinical settings.
The Food and Drug Administration has cleared over 1,400 AI-enabled medical devices as of early 2026. These tools read X-rays, flag lab results, and direct patient flow. But a Nature Medicine analysis found that nearly half - 226 of 521 approved devices - lacked data showing real-world performance. Another review of about 700 authorized devices reported that less than 4% included racial or ethnic data for their test subjects. The result is a regulatory system that often clears devices without proving they work safely across diverse patient populations.
The illusion of training accuracy
Clinical machine learning tools are typically tested by splitting data into two groups: one for training the model, and a separate "quiz" group to measure accuracy. When the model does well on the quiz group, developers and evaluators consider the numbers reliable. The problem is that the quiz group comes from the same population as the training data. If the model is then deployed on patients with different demographics, in a different hospital, or with different health profiles, the accuracy can drop sharply. The physician relying on the system has no way to know how much the difference matters for their own patient.
Those who suffer most from these performance drops are often the people already underserved by medical research. The pulse oximeter, a device clipped to the finger to estimate blood oxygen, was calibrated largely on lighter-skinned individuals. A 2020 New England Journal of Medicine study found that Black patients were nearly three times more likely than white patients to have critically low oxygen levels that the tool missed. These devices don't fail randomly; they fail where medicine has always failed, now with a reassuring number on the screen.
Regulatory gaps and demographic blind spots
FDA clearance of medical devices does not confirm clinical efficacy. Many approvals rely on the "predicate" method, comparing a new device to one already on the market. No clinical trials or head-to-head studies are required to show the new device is better. Post-market testing is sometimes proposed as a safety net, but it offers little real protection - it effectively uses the patient population as a test group. As AI becomes more integrated into clinical workflows, demand for transparency in AI for Healthcare is growing. Yet the current system leaves clinicians with devices that may perform well on paper and falter in the exam room.
What hospitals and regulators can do
Instead of banning AI, hospitals can require vendors to demonstrate performance on their own patient populations and subpopulations. The FDA can mandate demographic data in performance reports. Payers can link reimbursement to validated clinical outcomes. These steps do not reject a valuable tool - they hold software to the same standard applied to everything else in clinical practice. The worthwhile tools will pass a true validation study. The ones that only show results on paper are the ones that need to be caught first.
Why this matters for healthcare professionals
Clinicians should not assume that FDA clearance means an AI device has been proven effective for their patients. The fact that a sepsis model worked in a training data set does not tell you how it will perform on the floor. When evaluating any AI-based tool, ask for evidence of validation on a population that matches your own. The tools that lack that evidence are not fit for clinical decisions - no matter how many hospitals have already installed them.
Your membership also unlocks: