A Model Cannot Be Validated Until Someone Writes Down What It Is For
Validation in regulated work has never been a property of software in the abstract. It is a statement that a specific system, in a specific version, does a specific job within defined limits. AI programs break this in a quiet way. The team validates “the model” and never writes down the job. A deviation summarizer, a batch record reviewer, and a literature triage assistant can run on the same language model and carry very different risk. Only the job tells you how much evidence you need.
The draft EU GMP Annex 22 makes the point sharply. As summarized by European Pharmaceutical Review, it covers static models with fixed parameters and states that generative AI and large language models should not be used in critical GMP applications. They remain open for non-critical work, such as summarising deviation reports or searching procedures, when a qualified person reviews the output and keeps documented responsibility. Whether the exclusion survives finalization is still open, since the text remains a draft. The five requirements around it read like ordinary validation discipline: document the intended use and the full input range before testing, define the performance baseline before testing, keep test data apart from training data, record which features drive decisions, and run change control on every change to the model, system, or inputs.
The joint FDA and EMA principles point the same way from a different direction. They call for a clear context of use, a risk-based approach, strong data governance, and lifecycle management from early research through manufacturing and post-marketing safety. Two agencies and one draft annex ask the same first question: what is this system for, and how bad is it if it is wrong? The draft annex adds a second one that matters for buyers. It places responsibility for the validation evidence on the regulated company, whether the model was built in house or bought.
An AI system is validated for a job, and a job nobody has written down cannot be validated.
Six Validation Gaps in Life Sciences AI, Ranked by How Hard They Are to Defend
| Gap | What goes wrong | Severity |
|---|---|---|
| No written intended use or context of use | Testing starts without a defined task or input range, so no result can be tied to a claim about what the system does | Critical |
| Test data overlaps tuning data | Cases used to write prompts or pick a model are reused as proof, so the pass rate says little about unseen records | Critical |
| Reviewer sign-off with no evidence | An SOP names a human reviewer, but nothing shows what they checked, and approval rates drift toward 100 percent | High |
| Version changes without retesting | A vendor model update or a prompt edit goes live, and nobody can say which version produced an old output | High |
| Vendor package accepted as your evidence | Supplier validation documents are filed without being checked against your own intended use | Moderate |
| No inventory of AI features in SaaS tools | AI functions arrive inside quality or lab systems through an upgrade and never enter the validation register | Lower |
Not sure where your life sciences AI validation gaps are?
10decoders reviews how your AI workflows are scoped, tested, and monitored, from intended use through reviewer evidence. You get a ranked list of gaps and a scoped plan to close the highest-risk ones first.
Book a Free AI Assessment →Non-Critical Is a Classification You Have to Earn
The draft annex leaves room for language models in non-critical work, and most teams will use that room. The risk is classifying by feel. An assistant that only drafts a deviation summary feeds a human who signs it, so it sounds safe. But if that reviewer sees forty summaries a day and approves forty, the human control exists on paper only. The Pharmaceutical Technology commentary on Annex 22 makes a related argument: the qualified state belongs to the combination of a person, a system, a version, and a use case, and technical controls alone cannot absorb that accountability.
So the review step has to be designed like any other control. Show the reviewer the source record next to the summary. Log what they opened, how long they spent, and what they changed. Sample the reviewers whose acceptance rate never moves. Then test the AI step itself against closed records. Pull a set of past deviations with approved summaries, set the acceptance criteria before the run, and score invented facts and omitted facts separately. The bar from the draft annex is that the model performs at least as well as the process it replaces, so measure the manual process first.
Change control is where good pilots go wrong later. A vendor updates the underlying model, someone edits a prompt to fix a complaint, or a new site’s records enter the index. Each is a change to the validated state. Record the model, prompt, and index version on every output, define which changes rerun the test set, and name the person who decides whether revalidation is needed. The 63 percent of leaders who mandate human validation have the right instinct. The evidence trail is what turns it into something an auditor can read.
The Model Is Tested, the Job Is Not Defined
A team trials an assistant on a few real records and likes the output. There is no intended-use statement, no acceptance criteria, and no record of who reviewed what.
Intended Use, Test Set, and Criteria Exist
Each workflow has a signed intended-use statement and a classification. A closed-record test set runs against criteria set in advance, but reviewer evidence and change triggers are still informal.
Evidence Is Produced by the System
Versions and reviewer actions are logged automatically. Defined changes rerun the test set, and production samples feed new cases back in. Inspection questions are answered from records.
Checklist: Could You Defend This AI Workflow in an Inspection?
Pick one AI workflow that touches a regulated process and check each item. Anything you cannot show as a document, a log, or a named owner is a gap.
If a reviewer approves everything the assistant drafts, the review is a signature and not a control.
What to Do This Week
01 Write the intended-use statement for your highest-volume workflow
Choose the AI workflow with the most daily runs. In one page, state the task, the document types and languages it accepts, the output format, who uses the output, the decision it feeds, and what it must never be used for. Add the critical or non-critical classification and the reasoning behind it. Have the QA owner approve it. If the group cannot agree on the page, you have found the real validation gap.
02 Build a 40-case test set from closed records
Pull 40 closed deviations, batch records, or whichever record type the workflow handles, each with its approved human output. Include awkward cases: missing fields, records from two sites, a rare product. Write the acceptance criteria before you run anything, for example no invented facts and an agreed ceiling on omissions. Keep these cases out of every prompt and tuning session so they stay usable as evidence.
03 Log what reviewers actually do
Add logging for each review: which source records were opened, time spent, edits made, and the accept or reject decision. After a week, list reviewers with a 100 percent acceptance rate and read a sample of their approved items next to the sources. The goal is to learn whether the control works, so treat the result as process data and not as a verdict on individuals.
04 Stamp every output with versions and name your change triggers
Store the model identifier, a hash of the prompt, and the index version with each output. Then write a short list of changes that force a rerun of the test set: vendor model updates, prompt edits, new document sources, new sites. Ask each AI vendor in writing how and when they notify you of model changes. Without that notice, you cannot meet your own change control.
Let 10decoders Review Your Life Sciences AI Validation Before an Inspector Does
We review one AI workflow end to end: intended use and classification, test set and acceptance criteria, reviewer evidence, version logging, and change triggers. You leave with a ranked gap list and a plan to reach a controlled lifecycle.
