Generative AI tools have found a place in everyday work. More recently, people are starting to turn to AI to support documentation and development of medical devices. AI can help draft meeting notes, summarize literature searches, or clean up formatting. But there is a specific category of deliverable where the convenience of AI drafting creates real regulatory and patient-safety risk: the Biocompatibility Evaluation Plan (BEP) and Biocompatibility Evaluation Report (BER).
This isn't a caution against AI as a category. It's a case for understanding exactly where the boundary sits, and why crossing it undermines the purpose of these documents.
A BEP and BER under ISO 10993-1 are not fill-in-the-blank exercises or templates. They require someone to reason through:
The nature and duration of patient contact (surface, externally communicating, implant; limited, prolonged, or permanent)
Which endpoints are triggered by that contact category, and which can be scientifically justified as not applicable
Whether existing data — chemical characterization, prior testing on materially equivalent devices, literature, clinical history — is sufficient to close an endpoint without new testing
How manufacturing changes, sterilization method, or a new supplier might alter the biological risk profile of an otherwise "unchanged" device
Many parts of the BEP and BER are based on professional judgement from an individual who is trained to understand and assess risk of the specific device, its intended use, and its full manufacturing history. An AI model does not have the professional judgement, training and education nor does it have an ability to weigh the benefits and risks that exist between the data that exists across multiple files such as your device master record, your supplier's material certifications, your sterilization validation data, or the toxicological profile of the specific resin. It can produce a document that looks like a BEP or BER, with the right headings, right ISO citations, right vocabulary, but without having done any of the underlying research and evaluation. That's the core problem: the format is easy to imitate; the reasoning is not.
AI language models are well documented to generate plausible-sounding but incorrect citations, invented study data, and confidently wrong toxicological thresholds. In most business writing, a fabricated citation is embarrassing. In a BER, a fabricated or misapplied toxicological threshold, an invented "equivalent predicate" comparison, or a misstated permissible exposure limit becomes part of a regulatory submission. If it isn't caught, it can:
Delay clearance when a reviewer catches the error and issues an hold letter
Pass through review and result in a device on the market with an inadequately characterized biological safety profile
Create a traceability problem later, when someone tries to determine why a conclusion was reached and finds no real analytical trail behind it
Regulatory reviewers increasingly know what AI-generated regulatory text looks like — generic, structurally correct, but thin on device-specific reasoning and free of the small inconsistencies that come from someone actually being close to the data. That pattern itself draws scrutiny.
A defensible biocompatibility rationale — particularly a toxicological risk assessment (TRA) for chemical characterization data — depends on:
Standards that exist behind a pay wall and AI cannot access, without your intervention
Actual extractables and leachables data from your specific device or material
Margin of safety calculations against tolerable intake values, which require correct exposure duration, patient population, and route of exposure
A qualified toxicologist's interpretation of what those numbers mean in context
AI models cannot run these calculations against your real data because they don't have your real data unless you paste it in, and even then, they're not performing validated toxicological risk assessment, they're pattern-matching to text that resembles one. A margin-of-safety (MOS) calculation that looks right but wasn't performed by someone qualified to interpret exposure and toxicity data is not a safety net; it's a liability sitting quietly in your design development file.