Blog - Rook Quality Systems

Why AI Should Not Write Your Biocompatibility Evaluation Plan or Report

Written by Heather Ferguson | Sep 28, 2026, 2:23:09 PM

Why AI Should Not Write Your Biocompatibility Evaluation Plan or Report

Generative AI tools have found a place in everyday work. More recently, people are starting to turn to AI to support documentation and development of medical devices. AI can help draft meeting notes, summarize literature searches, or clean up formatting. But there is a specific category of deliverable where the convenience of AI drafting creates real regulatory and patient-safety risk: the Biocompatibility Evaluation Plan (BEP) and Biocompatibility Evaluation Report (BER). 

This isn't a caution against AI as a category. It's a case for understanding exactly where the boundary sits, and why crossing it undermines the purpose of these documents.


 

The BEP and BER documents are risk-based judgments from people with specific training, not just a template

A BEP and BER under ISO 10993-1 are not fill-in-the-blank exercises or templates. They require someone to reason through:

  • The nature and duration of patient contact (surface, externally communicating, implant; limited, prolonged, or permanent)

  • Which endpoints are triggered by that contact category, and which can be scientifically justified as not applicable

  • Whether existing data — chemical characterization, prior testing on materially equivalent devices, literature, clinical history — is sufficient to close an endpoint without new testing 

  • How manufacturing changes, sterilization method, or a new supplier might alter the biological risk profile of an otherwise "unchanged" device

Many parts of the BEP and BER are based on professional judgement from an individual who is trained to understand and assess risk of the specific device, its intended use, and its full manufacturing history. An AI model does not have the professional judgement, training and education nor does it have an ability to weigh the benefits and risks that exist between the data that exists across multiple files such as your device master record, your supplier's material certifications, your sterilization validation data, or the toxicological profile of the specific resin. It can produce a document that looks like a BEP or BER, with the right headings, right ISO citations, right vocabulary, but without having done any of the underlying research and evaluation. That's the core problem: the format is easy to imitate; the reasoning is not.

 

 

Fabrication risk is not a hypothetical here

AI language models are well documented to generate plausible-sounding but incorrect citations, invented study data, and confidently wrong toxicological thresholds. In most business writing, a fabricated citation is embarrassing. In a BER, a fabricated or misapplied toxicological threshold, an invented "equivalent predicate" comparison, or a misstated permissible exposure limit becomes part of a regulatory submission. If it isn't caught, it can:

  • Delay clearance when a reviewer catches the error and issues an hold letter 

  • Pass through review and result in a device on the market with an inadequately characterized biological safety profile 

  • Create a traceability problem later, when someone tries to determine why a conclusion was reached and finds no real analytical trail behind it

Regulatory reviewers increasingly know what AI-generated regulatory text looks like — generic, structurally correct, but thin on device-specific reasoning and free of the small inconsistencies that come from someone actually being close to the data. That pattern itself draws scrutiny.

 

 

Toxicological risk assessment requires access to documents AI doesn't have

A defensible biocompatibility rationale — particularly a toxicological risk assessment (TRA) for chemical characterization data — depends on:

  • Standards that exist behind a pay wall and AI cannot access, without your intervention 

  • Actual extractables and leachables data from your specific device or material 

  • Margin of safety calculations against tolerable intake values, which require correct exposure duration, patient population, and route of exposure 

  • A qualified toxicologist's interpretation of what those numbers mean in context

AI models cannot run these calculations against your real data because they don't have your real data unless you paste it in, and even then, they're not performing validated toxicological risk assessment, they're pattern-matching to text that resembles one. A margin-of-safety (MOS) calculation that looks right but wasn't performed by someone qualified to interpret exposure and toxicity data is not a safety net; it's a liability sitting quietly in your design development file.

 

Accountability doesn't transfer to a model

TFDA and Notified Body reviewers hold the Quality Representative, and leadership of a company accountable for the quality management (QMS) system, and the conclusions in a BEP and BER, not a tool. Design controls (21 CFR 820.30 / ISO 13485) require documented, traceable rationale connecting risk analysis to test data to conclusions, with sign-off from someone with the toxicological or biological safety expertise to own that judgment. 

If an AI-drafted rationale is submitted and later found to be flawed, "the AI generated it" is not a defense in an audit finding, a warning letter, or litigation. The responsible person under your quality system still owns that conclusion. Using AI to draft the substance of the document doesn't reduce that liability, it just obscures how thin the underlying analysis actually is.

 

Where AI can still add value in this workflow

None of this means AI has no place in the documentation of biocompatibility work. Used well, it can:
  • Summarize and organize literature search results for a human reviewer to evaluate 
  • Check formatting, internal consistency, or cross-references in a document someone else already wrote 
  • Draft boilerplate sections (document control headers, standard definitions) that carry no device-specific judgment
The line is consistent: AI can support the packaging of a BEP or BER. It should not generate the risk-based conclusions of the BEP and BER. That distinction of administrative support versus substantive regulatory judgment is the one worth holding onto as these tools get better and more tempting to lean on.

 

The bottom line

Biocompatibility Evaluation Plan and Report exist to demonstrate, with traceable rigor, that a device is biologically safe for its intended patient contact and duration. That rigor is the entire point. A document that reads correctly but wasn't built on genuine, device-specific analysis doesn't just risk a review delay, it risks releasing a device without the safety characterization the document claims to provide. Keep the drafting assistance. Keep the human, with real toxicological and regulatory expertise, doing the actual evaluation.
 
 If your team is working through a BEP or BER, we'd be glad to walk through it with you.