AI in GxP Systems: Your Validated System May No Longer Behave the Way You Think

Let’s be honest, there is no shortage of articles about AI in life sciences. Most of them say roughly the same thing: AI is here, and it’s going to change everything.

Let’s be honest, there is no shortage of articles about AI in life sciences.

Most of them say roughly the same thing: AI is here, and it’s going to change everything. OK, fair enough. But once you look past that, the more interesting question is much less glamorous:

What does this actually mean for validation in the systems we already rely on such as ERP, eQMS, LIMS, or Document Management?

Because AI isn’t a future discussion anymore. It’s already embedded in these platforms for classifying deviations, flagging anomalies, suggesting outcomes. Often in small, almost invisible ways. But increasingly in places that matter.

Same systems, different behaviour

On the surface, nothing has changed. The systems are still there, they’re still validated. GAMP guidance still applies as an effective and pragmatic framework. But the behaviour underneath is different.

Traditional systems are deterministic. You configure rules, test them, and expect consistency. AI changes that. Outputs are shaped by data, patterns, and probabilities. Not just by logic and familiar processes.

Which means:

  • The same input does not always produce the same output
  • Performance can shift over time
  • And perhaps most importantly: issues don’t always fail loudly

They drift.

That’s exactly where “traditional” CSV starts to feel uncomfortable. Not because it’s wrong, but because it was never designed for systems that behave this way.

A real-world example: when control is missing

In April 2026, the FDA issued a warning letter citing overreliance of AI in a GMP context. A company had used a generative model to draft drug specifications, procedures, and master production/control records.

The problem wasn’t the use of AI, it was how it was used.

The model smoothed inconsistencies that should have triggered human assessment, while outputs were insufficiently reviewed. In other words, AI started to replace (not support) critical thinking.

The FDA’s position was clear: using AI doesn't reduce a manufacturer's responsibility to understand and comply with its own regulatory obligations.

GAMP 5 still holds, if you apply it consciously

The takeaway from that example is not that we need a completely new framework, we really don’t. GAMP 5 still works, but it requires more deliberate application than many organisations are used to.

The fundamentals (intended use, risk-based thinking, lifecycle approach) are still exactly the right ones, but the bar is higher now.

Take intended use. In an AI context, it’s not just a description of functionality, it defines what you’re allowed to rely on. And in reality, that boundary tends to move. A suggestion becomes a recommendation, and a recommendation starts influencing decisions.

That shift rarely happens in a controlled way, and that’s where risk creeps in.

Risk is driven by impact, not by “AI”

Another trap we see quite often is treating “AI” as a category in itself. It’s not.

There is a big difference between AI suggesting metadata in a DMS, and AI influencing quality or manufacturing decisions through an ERP or eQMS. Same label, but with very different consequences.

What matters is simple: what happens if “it” gets it wrong?

Most current use cases sit somewhere in the middle, supporting human decisions, not replacing them. That human-in-the-loop principle is often the control that keeps things acceptable.

Remove it, and things go wrong very quickly, as the FDA example shows.

Validation becomes a true lifecycle

GAMP 5 already tells us validation is a lifecycle, especially in a world of continuously updated systems. AI doesn’t change that; it only makes it more visible.

System behaviour can change without a formal release, through evolving data or model configuration. So, maintaining the validated state requires continuous attention.

That puts monitoring front and centre: performance trends, drift, how users actually use the output. Not as an extra control, but as part of validation.

Turning principles into practice

None of this is conceptually new, most organisations already understand it. But translating it into something that works in real systems, within a real QMS, that’s where things get difficult.

At Aegis Consulting, we spend time helping organisations do exactly that. Not by reinventing validation, but by making GAMP 5 work for how systems behave today.

That usually comes down to being very clear about:

  • what the system is supposed to do
  • what you rely on
  • where the real risk sits
  • and how you prove you are in control over time

Simple in theory, harder in practice.

Final thought

AI is already part of your GxP landscape, acknowledged or not. And expectations around it are evolving quickly. The real question is not whether you’re using AI. It’s whether you can confidently explain (internally and to e.g. an inspector) why you trust it. If that’s not a straightforward answer yet, you’re not alone.

A note of credit

If you’re working on this topic, I’d strongly recommend the recently published handbook “Validating AI in GxP Environments” by Sachin Bhandari. It’s one of the more practical and grounded resources out there right now.