Healthcare: FDA, MDR, and the clinical validation gap
Lesson 1 of 5 in Sector by Sector: Health, Employment, Finance, Education, and Vehicles.
Healthcare is where AI governance is oldest — medical device law has regulated software since the 1990s — and where its central weakness shows most clearly: market authorisation is not clinical validation.
Start with the US pathway. The FDA has authorised over a thousand AI-enabled medical devices, concentrated in radiology. Almost all arrived via the 510(k) route: show substantial equivalence to an already-marketed predicate device — a pathway that typically requires no new prospective clinical trial. The result is a market full of legally authorised algorithms whose real-world performance across sites, scanners, and populations was never demonstrated pre-market.
The cautionary tale is not even an FDA-cleared product. The Epic sepsis model, embedded in the EHR used across hundreds of US hospitals, escaped device review as decision support. When University of Michigan researchers externally validated it (JAMA Internal Medicine, 2021), the advertised discrimination collapsed — it missed roughly two-thirds of sepsis cases while burying clinicians in alerts. Hundreds of hospitals had deployed it on the vendor’s internal numbers. External, site-level validation is the control this sector keeps re-learning.
The FDA’s genuinely novel contribution is the Predetermined Change Control Plan (PCCP), finalised in guidance in December 2024. Classic device law froze the approved artifact — every model update meant a new submission, which is fatal for systems designed to be retrained. A PCCP lets the manufacturer pre-specify what will change (retraining data, performance ranges), how changes will be validated, and the impact assessment — and then update within that envelope without returning to the agency. It is the first regulatory instrument built for drifting, learning systems, and other sectors are studying it closely.
In the EU, an AI diagnostic tool answers to two regulations at once. The Medical Device Regulation (MDR) classifies most diagnostic/therapeutic software as class IIa or higher (Rule 11), requiring notified-body conformity assessment. The AI Act then designates AI safety components of MDR-covered devices as high-risk via the Annex I route — Article 6(1) — layering its data-governance, logging, transparency, and human-oversight requirements on top of MDR clinical evaluation. Manufacturers run one integrated conformity assessment through their notified body covering both regimes; the practical friction (notified-body capacity, overlapping documentation) is the subject of ongoing Commission/MDCG guidance. Globally, the WHO issued governance guidance for AI in health (2021) and for large multimodal models in health (January 2024) — the first WHO guidance aimed at generative AI, warning against deploying LMMs in clinical workflows without task-specific evaluation.
US (FDA)
Instrument: device law — 510(k)/De Novo/PMA, PCCPs for model updates, post-market surveillance.
AI-specific moves: 1,000+ AI-enabled device authorisations; PCCP final guidance (Dec 2024); Jan 2025 draft guidance on AI-enabled device software functions covering the total product lifecycle; ongoing policy work on where clinical decision support stops being an exempt aid and becomes a regulated device.
Gap: wellness apps, administrative AI, and much EHR-embedded decision support (the Epic sepsis pattern) sit outside device review entirely.
EU (MDR + AI Act)
Instrument: MDR/IVDR conformity assessment (Rule 11 pushes most medical software to class IIa+), with the AI Act layered on via Article 6(1)/Annex I for AI safety components.
AI-specific moves: integrated notified-body assessment covering both regimes; AI Act adds dataset representativeness, logging, and human-oversight duties MDR never spelled out.
Gap: notified-body capacity is the chokepoint, and the double regime raises costs that fall hardest on clinical-AI startups.
Global (WHO)
Instrument: soft law. WHO Ethics and governance of AI for health (2021, six principles) and the LMM guidance (Jan 2024) with 40+ recommendations across the value chain.
AI-specific moves: the LMM guidance names concrete risks — automation bias in clinicians, degraded performance on under-represented populations, unvalidated diagnostic use of general-purpose chatbots.
Gap: no enforcement; its power is as a template for health ministries writing national rules, especially where no device regulator has AI capacity.
Key terms: samd, pccp, clinical validation, notified body, external validation
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.