FDA and generative AI: no new policy, but a regulatory framework is taking shape
On 18 August 2026 the FDA published its discussion paper on the regulation of generative AI-enabled medical devices. It is not draft guidance, not final guidance and not a new policy statement. It should not, however, be dismissed as merely exploratory.
Diana Hohage
Principal Consultant
In brief
FDA's discussion paper sets out a two-axis risk framework, a competency-based approach to premarket evidence and risk-proportionate postmarket monitoring for generative AI-enabled devices. Public comments are open under Docket FDA-2026-N-7874 until 19 October 2026.
On 18 August 2026, the U.S. Food and Drug Administration published its discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices.
The legal status of the document is important. It is not draft guidance, not final guidance, and not a new FDA policy statement. The Agency expressly states that the paper does not establish new regulatory expectations or evidentiary requirements for future marketing submissions. Nor does it resolve whether the approaches under discussion could be implemented under FDA's existing authorities or would require new ones.
That said, the paper should not be dismissed as merely exploratory.
It sets out a notably more structured consultation on how FDA may approach GenAI-enabled medical devices across the total product lifecycle: from risk characterisation and premarket evidence to postmarket performance and change control. This is not the first time FDA has addressed lifecycle issues for GenAI, but it is a more specific articulation of the questions now being put to stakeholders.
A preliminary point on scope is essential. FDA does not regulate GenAI as a technology category. It regulates products that meet the legal definition of a medical device. The regulatory analysis will therefore remain anchored in the product's intended use, its functionality, and the risks associated with relying on its outputs.
Three elements are particularly noteworthy.
1. A more nuanced framework for risk
FDA is considering a two-axis framework. The first axis concerns device activity: the extent to which a function provides information, directs a user towards an action, or takes action itself, including the degree and independence of that activity. The second axis concerns the consequences of relying on an incorrect output, including the potential severity of harm.
This is more nuanced than treating autonomy as the sole risk driver. An ostensibly informational function may become action-directing because of its wording, specificity, personalisation, or clinical context. The paper also asks how risk should be assessed in multi-turn interactions, in patient-facing use cases, and where care-escalation errors can cause harm through either under-escalation or over-escalation.
The practical implication is clear: a future evidence strategy may need to address not only a device's intended medical purpose, but also what the device does in practice, how independently it acts, who relies on its output, and the consequences of error.
2. A competency-based approach to premarket evidence
FDA discusses a potential premarket approach combining non-clinical device benchmarking with clinical confirmation.
The critical unit of assessment would be the final, user-facing device in its intended deployment configuration, not the foundation model in isolation. The depth and breadth of the assessment would be calibrated to the device's intended use and risk profile.
For relevant devices, FDA discusses benchmarking that extends beyond conventional test-set performance. Potential dimensions include safety-critical recognition and escalation, adherence to intended scope, calibrated communication of uncertainty, clinical reasoning, quantitative analysis, communication quality, robustness, reproducibility, and performance across clinically meaningful subgroups. For agentic systems, the paper additionally raises multi-step planning, reliable tool use, human-oversight checkpoints before high-consequence actions, and resilience to prompt injection.
The broader point is that even extensive non-clinical benchmarking may not fully establish how a GenAI-enabled device will perform in clinical practice. FDA therefore considers whether clinical confirmation under real-world or clinically representative conditions may be necessary. This would not necessarily require a prospective clinical study in every case; the required evidence would remain risk-proportionate.
Put differently, the emerging question is not simply whether a model can generate an acceptable response under predefined conditions. It is whether the deployed device can operate safely within its intended clinical role, recognise when it should defer or escalate, communicate uncertainty appropriately, and perform reliably across relevant users, settings, patient populations, and interaction patterns.
3. Postmarket performance may become integral to the evidence strategy
FDA also discusses potential approaches to risk-proportionate postmarket monitoring. These include periodic re-benchmarking, sample-based review of real-world inputs and outputs by qualified independent clinicians, and monitoring for performance degradation or drift arising from changes in the input population, data environment, underlying model, or deployment architecture.
This is particularly relevant where a device incorporates a third-party foundation model. The paper distinguishes among sponsor-initiated updates, retraining or adaptive changes, and unplanned changes introduced by the developer of an underlying third-party model. FDA is exploring whether a premarket competency assessment could provide the baseline against which the device is re-benchmarked after modification. It also discusses Predetermined Change Control Plans as one potential mechanism for certain changes and seeks feedback on voluntary Foundation Model Device Master Files. Importantly, any such master file would not transfer responsibility: the device sponsor would remain responsible for demonstrating the safety and effectiveness of its own device.
What should manufacturers do now?
Manufacturers should not treat this paper as a new compliance checklist. It is not one.
They should, however, treat it as a useful stress test for their development and evidence strategy. That means defining the intended clinical role and behavioural boundaries of each GenAI-enabled function; documenting when the system must defer, escalate, or require human review; evaluating the deployed device rather than relying on claims about an underlying model; and establishing a defensible baseline for future modifications and postmarket monitoring.
For products that depend on third-party models, the practical questions are increasingly concrete. What is version-controlled? Which supplier-initiated changes could affect safety or effectiveness? Which test assets and acceptance criteria support re-benchmarking? How will performance degradation, subgroup effects, unexpected interactions, and safety-critical failures be detected, investigated, and governed?
The policy is not settled. The direction of travel, however, is becoming clearer.
For GenAI-enabled medical devices, validating an algorithm at a single point in time may not be enough. FDA is exploring how manufacturers can demonstrate control over the behaviour of the deployed device throughout its lifecycle.
Public comments are open under Docket FDA-2026-N-7874 until 19 October 2026.
Prepared from FDA primary sources; this article is for informational purposes and does not constitute legal advice.
Relevant for your project?
Similar questions in your current project?
In a first call we clarify what is specifically relevant for your situation, without obligation.
Request a call →Life Science Journal
Regulatory updates, straight to your inbox.
New requirements, authority decisions and practice notes. Once a month, unsubscribe any time.
Regulations & standards considered
- FDA, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (18 August 2026)
- Docket FDA-2026-N-7874 (public comments until 19 October 2026)
- FDA, DHAC Executive Summary: Total Product Lifecycle Considerations for Generative AI-Enabled Devices (20 to 21 November 2024)
- Predetermined Change Control Plans (PCCP) for AI-enabled device software functions
FAQ
Frequently asked questions
Related expertise
EU AI Act →
Classification and duties for AI in regulated products, into which this discussion would fit
Software as a Medical Device →
Lifecycle, version control and change management for the deployed device
Post-Market Surveillance →
Where re-benchmarking, drift monitoring and sample-based output review would be anchored
Related projects
All case studies →Sources
- U.S. FDA, Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (18 August 2026)
- Regulations.gov, Docket FDA-2026-N-7874
- U.S. FDA, DHAC Executive Summary: Total Product Lifecycle Considerations for Generative AI-Enabled Devices (20 to 21 November 2024)
Related insights
All insights →Your project
Have a concrete project?
Briefly outline your situation. We'll respond with an initial assessment, usually within one business day.
Prefer direct? +39 02 8904 1000
info@theentourage.it
- Reply usually within one working day
- 4 offices: DE · CH · IT · US
- 100% life sciences




