Data integrity has a consistent and specific meaning in a GxP environment.
ALCOA is our starting point: data must be attributable, legible, contemporaneous, original, and accurate. And the extended interpretation also includes complete, consistent, enduring, available, and traceable (ALCOA++).
Data integrity regulations and guidelines are mostly variations on the ALCOA++ schematic. There is nuance, however. Understanding their commonalities and divergences is relevant when building a data governance program, responding to an inspection or audit finding, or mapping a validation program.
Distinguishing between regulations and guidelines
- Regulations are issued by a statutory authority (FDA, EMA, other national authorities). They are legally enforceable. For example, non-compliance with 21 CFR Part 11 and EU GMP Annex 11 carries legal consequences like warning letters and 483s.
- Guidelines (GAMP 5, PIC/S PI 041-1, MHRA’s GxP guidance, WHO Annex 5) are not law. They represent a consensus of how to meet the regulation. Compliance is technically voluntary, but inspectors use these documents as a functional checklist. Deviating will likely require a thorough, risk-based justification during an audit.
Regulations spell out the legal requirement. Guidelines explain the accepted path. A validation program must trace a guideline’s checklist back to the actual regulatory requirement it supports. Auditors want to know the “why,” not just the “what.”
How current guidelines define data integrity
The core language is consistent across sources, but the emphasis shifts depending on who wrote it.
| Source | Current version | What it emphasizes |
|---|---|---|
| GAMP 5 (ISPE) | 2nd Edition, 2022 | Data integrity is one of three equal objectives (patient safety, product quality, data integrity) of a risk-based, critical-thinking approach to computer system validation. It reinforces ALCOA+ throughout (terminology ISPE had been using since its 2017 GAMP Records and Data Integrity Guide), and added explicit cloud/SaaS validation and AI/ML lifecycle guidance. |
| PIC/S PI 041-1 | Final, in force 1 July 2021 | At 63 pages, this is the most exhaustive. It covers data governance, data lifecycle ownership, data criticality/risk assessment methodology, and inspector-facing expectations for both paper and computerized systems. Written for inspectors, it’s the industry’s go-to operational reference. |
| MHRA GxP Data Integrity Guidance and Definitions | Revision 1, March 2018 | Defines data integrity as the extent to which data is complete, consistent, and accurate throughout its lifecycle. It strongly emphasizes data criticality and risk assessment as the starting point to ensure a blanket, one-size-fits-all approach is not applied with the same level of control. |
| WHO Annex 5 | Guidance on Good Data and Record Management Practices, 2016 | Broader applicability across resource-constrained and hybrid paper/electronic environments; reinforces the same lifecycle principle with more attention to manual and semi-automated systems. |
| FDA Data Integrity and Compliance With Drug CGMP | Final guidance, December 2018 | This Q&A format addresses common findings of FDA investigators. Data integrity examples include shared login credentials, audit trail ambiguity, data exclusion practices, backdating. This is cited often in Form 483s for US operations. |
These guidelines are not in opposition. The core ideas of data integrity carry across all of them. If your data governance SOP encompasses more than one jurisdiction, auditors will expect to see references to their “local” guidance.
Static vs. dynamic data
Guidance from FDA and MHRA draw a distinction between “static data” and “dynamic data”:
- Static data is fixed once recorded (for example, a batch record) and controls prevent downstream alterations.
- Dynamic data represents data with which a user can interact, reprocess, requery, or recalculate (for example, raw chromatographic files, database records, most modern computer system outputs). Data integrity requirements extend to the complete underlying record of metadata and audit trails, not just the reviewed printout or report that only represents one moment in time.
The old practice of printing a chromatogram or spectrum, signing the document (either physically or digitally), and then discarding or ignoring the original electronic file is extremely non-compliant in the eyes of both agencies.
Retaining only static, human-readable outputs fails DI expectations. Any validation lifecycle for a computerized system must account for dynamic outputs and confirm the full dataset behind the record.
Assessing data criticality for control rigor
Not every data point in a GxP environment warrants the same level of control.
In fact, handling every data point with equal rigor is fundamentally inefficient and creates its own business risk by diluting the attention high-risk data receives.
PIC/S PI 041-1 provides the logic:
- Identify the GxP impact. Does the data directly support a quality decision, batch release, or patient safety determination? High-impact data demands tighter controls like robust audit trails, restricted edit rights, and independent verification.
- Assess inherent integrity risk. This depends on the process. For example, a manual transcription step conducted by a human has different inherent risk than an automated system with a validated audit trail and access controls.
- Apply proportionate controls. Impact and inherent risk combine into a control level. Applying critical rigor to a facilities log that has no product quality bearing makes no sense. Conversely, under-controlling a step just because it’s automated introduces blind spots.
- Document the rationale. The assessment (not just the control) is of high interest to inspectors. Documenting each data point’s control strategy demonstrates defensible critical thinking.
Where data integrity fails in inspections and audits
Data integrity issues are recurring finding categories across FDA 483s, MHRA inspection reports, and PIC/S audits. These are the most common:
- No defined data ownership. Without a named data steward, “who’s responsible for reviewing this audit trail” becomes an uncomfortable question. Data governance must specify ownership at the system level.
- Documentation that can’t reconstruct history. If a data point changed and there’s no traceable record of the who/when/why, the record is indefensible. This is a common cause of data integrity 483 citations.
- Validation and verification gaps. Data entering a system without checking accuracy at the point of entry introduces all kinds of downstream errors that might manifest in a batch record review.
- Weak access control and audit trail configuration. Shared logins, disabled or unreviewed audit trails, and unrestricted admin access still appear as findings. This is usually a system configuration and validation gap solvable with technical controls.
- Training that doesn’t stick. Data integrity training delivered during onboarding, with no periodic reinforcement or role-specific content, tends to surface as a root cause in CAPA investigations.
What data integrity means for a validation strategy
Data integrity must be designed into the system and the validation approach from the start.
Furthermore, risk-based logic applies to data itself. Not every system or requirement needs the same validation rigor, but every decision about rigor needs a documented rationale.
Static, page-based, or document-centric validation records inherit the integrity risks of static data: version control ambiguity, disconnected audit trails across “documents,” and no single source of truth for a given requirement across its lifecycle.
Object-based architectures such as VLMcare are structurally better positioned to keep dynamic data (with full audit trail and metadata) intact and reviewable across the requirement’s lifecycle, rather than collapsing it into a static, page-bound artifact at the point of sign-off.

