← Back to blog
NewsAbout 7 min read

The FDA Is Writing Rules for AI Devices Nobody Has Built Yet

Published Oct 7, 2026
The FDA Is Writing Rules for AI Devices Nobody Has Built Yet

Regulators usually react. A product ships, something goes wrong, and the guidance follows. The FDA's device center just did the opposite. On October 1 it published its proposed guidance agenda for fiscal 2027, and near the top of the list sits a draft guidance document for a category of device it has not cleared a single example of.

The category is generative AI-enabled conversational devices intended for mental health disorders. The FDA's Center for Devices and Radiological Health has authorized more than 1,200 AI-enabled medical devices to date. None of them uses generative AI for a mental health indication. The agency is writing evidentiary recommendations before the first application arrives, which is a way of setting the bar early instead of negotiating it case by case later.

That move sits on top of a discussion paper CDRH released in August, titled Considerations for the Regulation of Generative AI-Enabled Medical Devices. It is explicitly a discussion document, not policy, and it carries 26 questions for the public to answer. Comments close on October 19. Anyone building a clinical product on a large language model has roughly two weeks to say something about how they want to be regulated.

A two-axis risk framework

The paper's core idea is that generative AI devices fail in ways that older software rules did not anticipate. The agency proposes scoring risk on two axes. One is the consequence of an incorrect output, meaning how badly a wrong answer hurts the patient. The other is the level of independence, meaning how much the system acts on its own versus handing a human a suggestion.

That combination is familiar in spirit. Radiation software that flags a fracture for a radiologist carries a different risk profile from software that makes the diagnosis without review. What is new is applying it to systems that are non-deterministic, meaning the same input can produce different outputs, and that can keep adjusting after deployment. Traditional software validation assumes a fixed artifact. A model that drifts does not fit that assumption.

To handle it, the paper floats an approach modeled on how clinicians are credentialed. The argument is that doctors are not tested on every scenario they might face. They pass structured assessments of underlying knowledge and reasoning, then practice under supervision. The FDA asks whether a generative device could be evaluated the same way, with non-clinical benchmarking against defined standards followed by clinical confirmation in real or simulated environments.

Post-market surveillance gets the harder assignment. Because these systems can change, the paper suggests tracking performance drift, hallucinations and unexpected outputs from both the device's own model and any foundation model underneath it. That second part matters. A device maker that fine-tunes someone else's model inherits behavior it cannot fully see, and the paper asks directly how pre-market and post-market expectations should apply when the manufacturer does not control the underlying model.

What the paper avoids saying

Read the document for what it drops and the direction gets clearer. It never uses the FDA's old distinction between locked and adaptive algorithms. It avoids the terms software in a medical device and software as a medical device, which were the vocabulary of a decade of device review. Clinical decision support appears only in a footnote.

The omissions look deliberate. The old categories were built for software that behaves predictably once validated. Generative systems blur the line between the tool and the output, and the agency appears to be clearing space for a framework that does not inherit those assumptions. Whether that is a genuine reset or a repackaging of the same principles is the thing stakeholders are being asked to comment on.

The guidance agenda says where the money is going

The fiscal 2027 agenda, which takes comments through November 30, tells you which drafts the agency considers close to settled. Three sit at the top. First, finalizing the AI-enabled device software lifecycle guidance first drafted in January 2025, which covers how a model is monitored, updated and revalidated across its life rather than only at clearance. Second, finalizing the predetermined change control plan guidance from August 2024, which defines what a sponsor must pre-specify to make certain post-market algorithm changes without filing a new submission. Third, finalizing guidance for robotic-assisted surgical devices, including a separate draft issued in late September covering instrument control, latency, cybersecurity and sterilization.

Those three are the agency saying: here is where deployment already is. Predetermined change control plans are the mechanism that lets an approved AI feature keep learning without a fresh review, and they are the difference between a model that ships once and a model that improves. Lifecycle guidance is what turns that permission into an operating requirement. Together they describe a device that is regulated as an ongoing process rather than a sealed product.

The third-party model problem

One question in the paper deserves more attention than the others. It asks how pre-market and post-market expectations should apply when the manufacturer does not control the underlying foundation model. That situation is now the norm rather than the exception. A hospital system that wraps a general-purpose model for triage is not retraining it, and it has no access to the training data, the fine-tuning procedure or the change schedule.

The agency's own framing acknowledges the difficulty. Post-market surveillance, in its model, would have to track drift in the device's model and in the foundation model beneath it, which means a device maker needs visibility into changes it did not make and may not be told about. Extending the competency idea to this case raises the question of who gets credentialed. The device maker can demonstrate that its own layer behaves, but it cannot easily demonstrate the behavior of a component it licenses.

There is a comparison worth making with the earlier generation of AI device rules. Predetermined change control plans work when a sponsor controls the model and can pre-specify the changes it intends to make. When the model comes from a third party, the plan has to accommodate updates that arrive on someone else's schedule. None of the discussion questions resolve that, and the answers will largely determine whether the eventual framework is workable for small developers or only for firms large enough to build their own models.

Why the timing is unusual

Writing guidance ahead of a first application is rare, and the reason it is happening here is that the mental health conversational device is the hardest version of the problem. These systems talk to patients in natural language, which makes their outputs hard to bound and hard to test. A diagnostic model returns a number. A conversational one returns whatever it decides to say next.

The agency also has to reconcile its own caution with political pressure to speed up medical innovation, and the discussion paper reflects both. It leans on a risk-based, least-burdensome approach while acknowledging that pre-market review alone cannot evaluate a system that changes. That tension is why the document ends in questions rather than answers. It is also why the 26 questions are worth reading closely. They are a preview of the arguments that will shape the first approved generative mental health device, whenever someone files for one.

Other regulators are moving in parallel. The UK's medicines regulator sent three amendments to parliament to modernize device oversight, including explicit room for software and AI devices under a risk-tiered licensing framework. Japan is revising its biological product requirements to absorb newer molecular testing methods. The pattern across all three is the same: the rules are being rebuilt before the products arrive, on the bet that writing them in advance is cheaper than litigating them afterward.

Related articles