The Scenario:
A mid-sized CRO closes a routine GCP inspection with eleven findings, two of them critical: an EDC audit trail that had been disabled during a mid-study migration, and a pattern of unblinded data access that no one had reviewed for eight months. The response is due in thirty days. The quality director is short two headcounts, the data management lead is covering three studies, and the person with the deepest knowledge of the migration left the company in March.
So, the team does what a lot of teams are now doing. They feed the inspection report into a large language model, prompt it for a response, and get back a fluent, well-organized draft in under an hour. It looks professional. It cites regulations. It proposes CAPAs. A senior manager skims it, corrects two typographical errors, removes the em dashes so it does not look like AI created it, and submits it.
The response runs to ninety pages. It cites two MHRA guidance documents that do not exist. It quotes a regulatory framework that does not apply to the CRO’s activities. And it never actually addresses the audit trail deficiency, because the model had no organizational knowledge of what happened during the migration, and no one gave it any. It filled the gap with plausible language instead.
What the CRO submitted was not a response. It was a document shaped like one.
Artificial intelligence is a junior partner: fast, tireless, articulate, and entirely incapable of being held accountable. Accountability stays where it always was, with the human who has the ability to be rational and is charge.
What the Regulator Saw
This composite is built from a pattern the MHRA has now described publicly in a recent blog post. The Agency has received inspection responses generated with AI assistance that contained references to MHRA guidance that does not exist, citations of regulatory frameworks that were not relevant, and responses to serious deficiencies that appeared designed to obscure rather than resolve the underlying problem. According to the MHRA blogpost, one such response exceeded ninety pages and failed to address the deficiencies identified at all. In at least one case involving a deficiency with patient safety impact, an AI-generated response containing fabricated references and material inaccuracies delayed resolution of the compliance failure; review time rose from roughly four hours to more than twenty, and required a multidisciplinary team and a full reconciliation against previously issued guidance.¹
The Agency’s framing of the problem is the part practitioners should internalize. The MHRA states that the misuse of AI in this context has moved from theoretical risk to realized risk, and that its concern is not whether an organization uses AI, but whether submissions are accurate, verifiable, and prepared under appropriate oversight.¹ Nothing in the regulatory framework changed. Organizations have always been responsible for the accuracy of what they submit. Technical review has always been a core element of the quality system. Materially false statements to inspectors have always been an offence. Those obligations do not soften because the drafting tool improved.¹
Read that sequence carefully. The CRO in this scenario did not just write a bad document. It converted a manageable inspection outcome into evidence of a defective quality system, and it did so while trying to save time.
Examining the AI Value
None of this argues against using AI. The MHRA is explicit that the technology offers genuine benefits: articulating complex technical issues, improving consistency, accelerating routine drafting, and enabling innovation, and that used properly it can support better regulatory outcomes and patient safety.¹ The distinction is where in the workflow the tool sits.
Used well, AI compresses the mechanical work. It assembles the finding log, cross-maps each deficiency to the relevant sections of the protocol, the SOP, and the data management plan, drafts the structural scaffolding of the response, checks that every finding has a corresponding CAPA entry, and flags internal inconsistencies in terminology across a hundred pages of attachments. That work is real, it is tedious, and it consumes the hours of exactly the people who should be spending those hours somewhere else.
Somewhere else is the root cause. The MHRA is direct on this point: AI tools may support elements of the CAPA process but cannot substitute for the technical understanding and organizational knowledge required to develop meaningful corrective actions, and superficial or generic CAPA without evidence of genuine root cause analysis will be identified as a quality system weakness regardless of how the response was drafted.¹ The model does not know why the audit trail was disabled during the migration. It does not know that the same configuration decision was made on two other studies. It does not know that the change control record was signed by someone without the authority to sign it. A subject matter expert knows those things or can find them out. Give that expert the hours the model saved, and the response gets better. Skip the expert, and the model will invent something to fill the silence.
That is the entire argument. The value of AI in regulatory response is not that it replaces the SME. It is that it buys the SME time to do the work only the SME can do.
The Controls Needed
The MHRA’s stated expectations for any submission are that it be factually accurate and verifiable, technically reviewed by appropriately experienced people, signed off by someone with authority and accountability, supported by evidence for factual claims, and appropriate to the specific regulatory context.¹ Those five requirements are the acceptance criteria. Everything below is the control framework that makes them achievable.
The EU’s draft Annex 22 to the GMP Guide, published for stakeholder consultation on 7 July 2025 and still in finalization as of this writing, supplies the operational vocabulary.² It is written for manufacturing, and it is not yet an operative instrument. Its control architecture is nonetheless the most fully articulated regulatory statement available on how AI is to be governed in a GxP environment, and the principles transfer to clinical use without strain. Annex 22 requires a defined intended use, established performance metrics, controls on the quality of training data, management of test data, and continuous oversight including change control, model performance monitoring, and defined procedures for human review.³ The draft confines generative models to non-critical functions and requires documented human oversight at each decision point, reserving critical applications for models that behave deterministically.⁴
Translated into a clinical quality management system, the control set looks like this:

Map these to ALCOA+ and the logic is familiar. Accountable sign-off and AI-use disclosure protect attributability. Citation verification and SME review protect accuracy. Grounding the model in controlled source documents protects originality, in the sense that the assertion traces to a real record rather than to a statistical guess. Change control over model and prompt versions protects consistency. These are not new principles wearing new clothes. They are the same principles applied to a new author.
A Self-Audit Checklist Before Submission
The MHRA also published its warning signs, which function as a self-audit checklist before submission. Inadequate verification announces itself through factually incorrect statements or non-existent references, generic language inappropriate to the specific circumstances, absence of organization-specific detail where detail is expected, citation of regulatory frameworks without explanation of their relevance, inconsistent technical terminology across submissions, and verbose responses that fail to engage the subject matter.¹
Every one of those is detectable by a competent SME in a single careful read. That read is the whole control. A CRO that spends four hours on it protects an inspection outcome. A CRO that skips it spends twenty hours of a regulator’s time and buys itself a place on the risk register.
Sources
Brown, Peter. “Use of AI for GXP Inspection Responses: Setting Standards Without Stifling Innovation.” MHRA Inspectorate Blog, Medicines and Healthcare products Regulatory Agency, 29 June 2026, https://mhrainspectorate.blog.gov.uk/2026/06/29/use-of-ai-for-gxp-inspection-responses-setting-standards-without-stifling-innovation/.
European Commission. Stakeholders’ Consultation on EudraLex Volume 4, Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and New Annex 22. Directorate-General for Health and Food Safety, 7 July 2025, https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en. Annex 22 remains in draft and has not been adopted as an operative instrument; it is cited here for its control architecture, not as a binding requirement. EMA convened a multistakeholder expert workshop on 30 June and 1 July 2026 to inform finalisation.
European Commission. Draft Annex 22: Artificial Intelligence, EudraLex Volume 4, consultation guideline, July 2025, https://health.ec.europa.eu/consultations/stakeholders-consultation-eudralex-volume-4-good-manufacturing-practice-guidelines-chapter-4-annex_en.
European Medicines Agency. Multistakeholder Workshop on Expert Contributions to Artificial Intelligence Guidance Development (Annex 22), 2026, https://www.ema.europa.eu/en/events/good-manufacturing-practice-multistakeholder-workshop-expert-contributions-artificial-intelligence-guidance-development-annex-22. The draft’s confinement of generative models to non-critical applications, and its preference for deterministic behaviour in critical applications, are drawn from the consultation text and remain subject to revision.
International Council for Harmonisation. Guideline for Good Clinical Practice E6(R3), Step 4 version, adopted 6 January 2025. Sponsor quality management obligations at Sections 3.6, 3.9, and 3.10; sponsor accountability for delegated activities at Sections 3.6.5 and 3.6.7; system validation at Section 4.3.4.
The opening scenario is a composite illustration. It is not drawn from a single named organization and no identifying detail is intended. The regulator-side facts within it, including the ninety-page response and the escalation of review time from approximately four hours to more than twenty, are reported by the MHRA in source 1.
Check Out What Our Experts Have to Say
about trial design, data capture, operational efficiencies, and, ultimately, how to solve with certainty in clinical research.