← Back to all articles
Insights

Article 10 of the EU AI Act: A Plain-English Guide to Data Governance for High-Risk AI

Generated image

Most compliance teams preparing for the EU AI Act have focused on risk management systems, conformity assessments, and CE marking. Data governance - Article 10 - tends to get treated as a technical afterthought. That is a mistake.

For high-risk AI systems that are trained with data, Article 10 is arguably the most operationally demanding provision in the entire regulation. It reaches back to the earliest stages of model development, imposes documented governance requirements on every dataset you use, and creates a specific - and carefully bounded - permission to process sensitive personal data for bias correction. Getting it wrong is not a documentation gap. It is a substantive legal failure.

This guide explains what Article 10 actually requires, how it maps to your existing data and compliance workflows, and what you need to have in place before the enforcement clock runs out.


What Article 10 Is - and Who It Applies To

Article 10 of Regulation (EU) 2024/1689 (the EU AI Act) sets out mandatory data governance requirements for providers of high-risk AI systems that use techniques involving the training of AI models with data. If your system is trained - rather than purely rule-based - Article 10 applies to your training, validation, and testing datasets whenever those datasets are used.

Article 10 applies to high-risk AI systems that make use of techniques involving the training of AI models with data, and requires that training, validation, and testing datasets meet quality criteria set out in paragraphs 2 to 5 of the Article.

For high-risk AI systems that do not use training techniques - think expert systems or deterministic rule engines - paragraphs 2 to 5 apply only to testing datasets. That is a narrower obligation, but it is not zero.

The high-risk classification guide covers which systems fall into the high-risk category in detail. In short: Annex III systems (recruitment tools, credit scoring, biometric identification, critical infrastructure management, law enforcement, education, and others) and systems embedded in regulated products under Annex I Union harmonisation legislation both carry the full suite of Article 10 obligations.

star Important

Timeline update — Digital Omnibus (May 2026). On 7 May 2026, EU institutions reached a provisional political agreement deferring Annex III high-risk obligations to 2 December 2027, and Annex I product-embedded obligations to 2 August 2028. Formal adoption and publication in the Official Journal are expected before 2 August 2026. Until that formal publication, the original dates remain legally binding — and the substantive obligations are unchanged. Plan to the new dates, but do not pause the work.


The Five Things Article 10 Actually Requires

1. Data governance and management practices (Art. 10(2))

This is the heart of the Article. Providers must establish and maintain documented governance and management practices covering each of the following:

  • Design choices - what the data is intended to measure or represent, and the assumptions baked into that framing
  • Data collection processes and origin - where the data came from, how it was collected, and what selection criteria were applied
  • Data preparation operations - annotation, labelling, cleaning, updating, enrichment, and aggregation steps, including quality controls applied at each stage
  • Formulation of assumptions - explicit documentation of what the data is assumed to represent about the real world
  • Assessment of availability, quantity, and suitability - whether you have enough of the right data for the intended purpose
  • Examination for possible biases - a structured assessment of biases likely to affect health, safety, or fundamental rights, or lead to discrimination
  • Measures to detect, prevent, and mitigate biases - not just identification, but documented remediation
  • Identification of data gaps and shortcomings - known limitations of the dataset, and how they are managed

The regulation does not prescribe specific statistical tests or bias metrics. What it requires is that you have a documented, repeatable process - not a one-time pre-launch check. If you retrain your model on new data, you repeat the analysis for the updated dataset.

2. Dataset quality criteria (Art. 10(3))

Training, validation, and testing datasets must be:

  • Relevant - appropriate for the intended purpose
  • Sufficiently representative - covering the population and use cases the system will be applied to
  • Free of errors and complete - to the best extent possible, given the intended purpose
  • Statistically appropriate - including as regards the persons or groups of persons on whom the system will be used

The phrase "to the best extent possible" matters. The law does not demand perfection. It demands that you have genuinely tried, and that you can show your work.

3. Contextual appropriateness (Art. 10(4))

Datasets must take account of the geographical, contextual, behavioural, or functional setting in which the system will be used. A model trained on data from one jurisdiction, demographic, or operational environment cannot simply be deployed in a different one without reassessing representativeness. Treating a "one-size-fits-all" dataset as universally valid is a specific compliance risk under this paragraph.

4. Non-trained systems (Art. 10(6))

For high-risk AI systems that do not use training techniques, paragraphs 2 to 5 apply only to testing datasets. This is a narrower obligation but still requires documented governance over the data used to validate system performance.

5. Special-category data for bias detection (Art. 10(5))

This is the provision that surprises most teams - and it is genuinely useful if handled correctly.

Article 10(5) of the EU AI Act permits providers of high-risk AI systems to exceptionally process special categories of personal data (as defined in Article 9(1) GDPR) for the purpose of bias detection and correction, subject to strict conditions.

The conditions are cumulative - all must be met:

  • Bias detection cannot be effectively fulfilled by processing other data, including synthetic or anonymised data
  • The special-category data is subject to technical limitations on re-use, and state-of-the-art security and privacy-preserving measures, including pseudonymisation
  • The data is subject to measures to ensure it is not accessible to third parties
  • The data is not transmitted, transferred, or otherwise accessed by any third party
  • The data is deleted once the bias has been corrected, or when the retention period ends - whichever comes first
  • Records of processing activities include the reasons why processing special-category data was strictly necessary, and why the objective could not be achieved with other data

This is a deliberate carve-out designed to enable thorough bias auditing. But it is an exception, not a loophole. Regulators will scrutinise any use of sensitive data under this provision closely - and improper use risks exposure under both the AI Act and the GDPR simultaneously.

lightbulb Tip

Practical note on Art. 10(5). Before invoking this exception, document why synthetic or anonymised data is insufficient for your bias assessment. If you can achieve the same result with less intrusive data, you must use it. The burden of justification sits with the provider.


How Article 10 Connects to the Rest of the Regulation

Article 10 does not sit in isolation. It feeds directly into several other high-risk obligations:

Annex IV technical documentation. Your data governance records - dataset specifications, bias examination results, mitigation measures, data lineage - must be captured in the technical documentation required under Article 11 and Annex IV. Regulators assessing conformity will look for this documentation as evidence of diligence.

Quality Management System (Art. 17). Your QMS must cover data governance as a documented process, not just a one-off exercise. The first CEN-CENELEC harmonised standard to enter public enquiry - prEN 18286, covering quality management systems for AI Act regulatory purposes - specifically addresses this lifecycle governance requirement.

Conformity assessment (Art. 43). Data governance records form part of the evidence base for conformity assessment, whether self-assessed or third-party. The conformity assessment and CE marking guide explains how this documentation feeds into the broader conformity process.

Risk management (Art. 9). Bias risks identified under Article 10 must feed into the risk management system. The two obligations are designed to work together: Article 10 identifies the data-level risks; Article 9 manages them at the system level.


The Harmonised Standards Picture

CEN and CENELEC, working through their Joint Technical Committee JTC 21, are developing harmonised standards that will cover data governance and dataset quality for high-risk AI systems. Once published in the Official Journal of the EU, following those standards will give providers a presumption of conformity with the corresponding AI Act requirements - making compliance significantly easier to demonstrate.

CEN-CENELEC JTC 21 is developing harmonised standards covering datasets and bias as part of its AI Act work programme, with standards on datasets, bias, and quality management among the specific deliverables.

The standards are running behind the original schedule. The first standard - prEN 18286 on quality management systems - entered public enquiry in October 2025 but is not expected to be finalised and cited in the OJEU until late 2026 at the earliest. Dataset-specific standards are at earlier stages.

The practical implication: you cannot wait for the standards to tell you what to do. You need to interpret the legal text directly, document your approach, and be ready to demonstrate that your data governance practices satisfy the Article 10 requirements on their own terms. When harmonised standards do arrive, they will provide a clearer compliance pathway - but they will not change the underlying obligations.


Penalties for Non-Compliance

Breaches of provider obligations under Article 10 fall under the mid-tier penalty band in Article 99: up to €15 million or 3% of total worldwide annual turnover for the preceding financial year, whichever is higher.

For SMEs and start-ups, the lower of the fixed amount or the percentage applies. The Article 99 fines guide covers the full penalty structure, including how national market surveillance authorities will exercise their enforcement powers.

One important point: Article 99(8) prevents double penalties for the same factual violation under both the AI Act and the GDPR. But that protection only applies where the facts genuinely overlap - it does not insulate you from parallel GDPR enforcement for separate data protection failures.


Use This Tool: Article 10 Readiness Self-Assessment

Before you build your compliance programme, it helps to know where your gaps actually are. Use the interactive checklist below to assess your current Article 10 readiness across the five core obligation areas.


What to Do Before the Enforcement Deadline

The obligations are fixed. The deadline has moved - but the work has not. Here is a concrete checklist for providers of Annex III high-risk systems working toward the December 2027 enforcement date (and for Annex I systems, August 2028).

1
Inventory your datasets

Map every dataset used in training, validation, and testing for each high-risk AI system. Record the dataset name, purpose, usage phase (training/validation/test), data owner, and source. This inventory is the foundation for everything else.

2
Document data provenance and lineage

For each dataset, record where the data came from, how it was collected, what selection criteria were applied, and what preprocessing steps were performed (cleaning, normalisation, annotation, aggregation). The EU AI Act treats labelling methodology as a core documentation requirement — not a technical detail.

3
Define fitness-for-purpose and representativeness criteria

Write down, explicitly, what the data is intended to measure or represent. Assess whether the dataset covers the geographical, contextual, behavioural, and functional settings in which the system will be deployed. Document known gaps and how they are managed.

4
Run structured bias testing and record mitigations

Conduct a documented bias examination for each dataset. Identify biases likely to affect health, safety, or fundamental rights, or lead to discrimination. Record the testing methodology, findings, and the specific mitigation measures applied. This is not a checkbox — it must be repeatable and updated if you retrain.

5
Establish a data governance record in your Annex IV technical documentation

Capture all of the above in your technical documentation (Annex IV) and Quality Management System (Article 17). Assign data owners, define approval workflows for dataset changes, and establish escalation paths when quality or bias risks are identified. Governance frameworks that exist only as policy documents are unlikely to satisfy enforcement expectations.

6
Establish a lawful basis and safeguards if using special-category data for bias correction

If you intend to invoke Article 10(5), document why synthetic or anonymised data is insufficient, apply pseudonymisation and access controls, restrict re-use, and establish a deletion schedule tied to bias correction milestones. Update your records of processing activities to include the justification for strict necessity.

7
Build ongoing monitoring into your QMS

Article 10 compliance is not a one-time pre-launch exercise. If you retrain your model, you repeat the analysis. If you deploy into a new geographic or operational context, you reassess representativeness. Build these triggers into your post-market monitoring plan and QMS review cycle.


Frequently Asked Questions

help_outlineDoes Article 10 apply if we use a third-party foundation model and fine-tune it?expand_more

Yes, if you are the provider of the high-risk AI system, Article 10 obligations apply to you — including for any fine-tuning datasets you use. You should also seek transparency from the foundation model provider about the governance of their pre-training data, as this affects your overall risk picture and technical documentation.

help_outlineWhat counts as 'sufficiently representative' under Article 10(3)?expand_more

The regulation does not define a specific statistical threshold. 'Sufficiently representative' means the dataset covers the population and use cases the system will actually be applied to, with appropriate statistical properties for those persons or groups. You need to document your representativeness assessment and the criteria you applied — not just assert that the data is representative.

help_outlineCan we use synthetic data to satisfy Article 10?expand_more

Synthetic data can contribute to meeting Article 10 requirements, but it does not automatically satisfy them. You still need to document the synthetic data generation methodology, assess whether it is representative of the real-world population, and examine it for biases. Synthetic data is also specifically mentioned in Article 10(5) as a first resort before invoking the special-category data exception.

help_outlineHow does Article 10 interact with GDPR?expand_more

Article 10 obligations apply regardless of whether personal data is involved — they are AI-specific requirements, not data protection requirements. Where your training data does include personal data, GDPR obligations (purpose limitation, data minimisation, accuracy, security) apply in parallel. The Article 10(5) special-category data exception operates alongside GDPR Article 9, not instead of it — you still need a valid GDPR legal basis for the processing.

help_outlineDoes the Digital Omnibus delay mean we can pause Article 10 work?expand_more

No. The provisional agreement reached on 7 May 2026 defers the enforcement date for Annex III systems to 2 December 2027 — but the substantive obligations are unchanged. The delay gives you more time to implement, not a reason to defer starting. Technical documentation from scratch typically takes 3–6 months, and bias examination requires access to training data that may no longer be readily available if you wait.


Where to Go Next

Article 10 is one piece of a larger compliance picture for high-risk AI providers. The data governance work you do here feeds directly into your risk management system (Article 9), your technical documentation (Article 11 and Annex IV), your QMS (Article 17), and your conformity assessment (Article 43).

If you have not yet confirmed whether your AI systems are high-risk, start with the Risk-Tier Classifier - answer a short questionnaire and get a provisional tier assessment with a plain-English rationale. Once you know your tier, the Obligations Checker maps every relevant Article 10 and broader provider obligation to practical compliance steps.

For the latest regulatory developments - including Digital Omnibus updates, harmonised standards progress, and AI Office guidance - subscribe to The AI Act Brief, our free weekly newsletter.


This guide is for informational purposes only and does not constitute legal advice. The EU AI Act is a complex regulation, its supporting standards are still being developed, and the Digital Omnibus amendments are subject to formal adoption. Always verify your obligations against the official text of Regulation (EU) 2024/1689 and the latest guidance from the EU AI Office. Consult qualified legal counsel for advice specific to your organisation.