FDA proposes competency-based evaluations for generative AI medical devices

The FDA is proposing a shift to competency-based testing for generative AI medical devices, with public feedback due by Oct. 19, 2026. The agency has already authorized over 1,000 AI-enabled devices, but the new framework would require benchmark suites, postmarket monitoring, and contract updates...

Categorized in: AI News Healthcare
Published on: Aug 30, 2026
FDA proposes competency-based evaluations for generative AI medical devices

The FDA has set an Oct. 19, 2026 deadline for public feedback on a discussion paper that proposes shifting how generative AI medical devices are evaluated - from feature-based claims to competency-based testing with heavier postmarket monitoring. For hospitals, clinics, and medical device buyers, the paper signals that the evidence behind genAI tools will soon need to look like a test plan, not a demo.

The agency's Center for Devices and Radiological Health has already authorized more than 1,000 AI-enabled devices, but most are not generative AI, according to Healthcare Dive. The discussion paper aims to draw a clearer line before genAI becomes routine in clinical documentation, imaging, and patient-facing guidance.

FDA proposes a competency-based approach

The FDA's paper outlines a "competency-based approach" to premarket evaluation, drawing from how human clinicians are assessed and credentialed, Healthcare Dive reported. The model combines nonclinical benchmarking with clinical confirmation.

Benchmarking is meant to measure clinical proficiency and generalizability, and the FDA explicitly calls out "agentic" capabilities in its discussion, according to Healio. That forces product teams to translate model behavior into measurable tasks and thresholds that can be reproduced outside the lab.

Clinical confirmation could take several routes. Healio reported the FDA is considering retrospective analysis of real patient data and "shadow deployment" in live workflows that doesn't affect actual care. Many health systems already use parallel runs, silent modes, and staged rollouts to validate new software. The shift is the FDA's willingness to accept those methods as part of formal evidence, not just local practice.

Postmarket monitoring becomes part of product design

The FDA is asking whether it should accept greater premarket uncertainty in exchange for greater reliance on postmarket monitoring, Healio reported. Medical Economics also identified risk-proportionate postmarket monitoring as a central theme of the paper.

That shift changes staffing and system requirements. Manufacturers will need instrumented products that can detect drift, measure field performance, and explain model updates. It also brings real-world performance into contract talks with health systems, since monitoring depends on data access, logging, and agreement on what counts as a reportable issue.

Healthcare Dive noted the scale of the challenge: the FDA's Center for Devices and Radiological Health has authorized more than 1,000 AI technologies, but most follow older AI playbooks. "The paper aims to draw that line before genAI becomes common in clinical documentation, imaging workflows, and patient-facing guidance," Healthcare Dive reported.

Healthcare professionals in administrative roles will see shifts in how AI tools are validated. For those working in medical billing or records management, the practical effect is that vendors will need to document what their tools do and how they perform over time, which can lead to clearer training and more reliable daily use. That's directly relevant for anyone working with AI for Medical Billers or AI for Medical Records Clerks, where tracking AI behavior matters as much as using it.

What this means for vendor selection and contracts

This is not new policy yet. The paper is open for discussion and early input, as Healio and Healthcare Dive both point out. But the direction is clear enough that buyers planning roadmaps for 2027 can use it now as a procurement filter.

The near-term work sits in three areas: evidence packaging, monitoring design, and change control. Competency language should push vendors to publish benchmark suites and limits, and it gives provider IT and clinical engineering teams a clearer way to ask, "What does this do, exactly, and how do we know when it stops doing it?"

Generative AI features built on foundation models raise a thornier governance issue. Medical Economics said the FDA is explicitly raising questions about foundation models and more agentic systems. Procurement and legal teams should tighten contract language now around model updates, retraining, and who approves a change when it could alter clinical behavior.

Where to begin pressure-testing a genAI device plan before Oct. 19

For regulatory and quality assurance teams: Can each genAI claim be mapped to a stated competency with a defined test set, acceptance criteria, and failure modes that can be audited later - including after model updates?

For product and clinical teams: Which clinical confirmation route fits for the intended use, such as a prospective study, retrospective analysis, shadow deployment, or clinician review? What data approvals are needed, and what timeline would it take to get them?

For health systems and procurement: Does the vendor's contract explicitly grant the logging, performance data access, and update controls needed to support ongoing monitoring? The FDA expects more postmarket evidence, and without data access, no monitoring program gets off the ground.

For both sides: If the genAI component relies on a foundation model, what is the update cadence? Which changes trigger revalidation? And how will "agentic" behaviors be constrained and tested? Medical Economics reported both foundation models and agentic systems are in scope for the discussion paper.

Why this matters for healthcare professionals

The fastest path to adoption is the unglamorous one: explicit competencies, explicit thresholds, explicit monitoring, and explicit update rights. Healthcare professionals who work with AI tools should expect a shift from flashy demos to rigorous, continuous evaluation. That will translate into more documentation and more structured training around AI systems - and in turn, more clarity about what these tools can reliably do in day-to-day clinical operations.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)