Art 10: data and data governance
Lesson 2 of 6 in Inside the High-Risk Rulebook: Arts 8–15 Article by Article.
You learned in Foundations that a model’s behaviour is grown from its data. Art 10 is the legislator drawing the obvious conclusion: if the data is where bias enters, the data is where the law must look. It attaches duties to all three data roles — training, validation and testing sets — and works on two levels.
Level one: governance practices (Art 10(2)). The datasets must be subject to data governance and management practices appropriate to the intended purpose, covering the full pipeline: relevant design choices; collection processes and origin of the data (and, for personal data, its original purpose); preparation operations — annotation, labelling, cleaning, updating, enrichment, aggregation; the formulation of assumptions about what the data measures; an assessment of availability, quantity and suitability; an examination for possible biases likely to affect health, safety or fundamental rights or lead to prohibited discrimination — with special attention to feedback loops, where the system’s outputs contaminate its future inputs; measures to detect, prevent and mitigate those biases; and identification of gaps and shortcomings plus how they will be addressed.
Level two: dataset quality criteria (Art 10(3)–(4)). The sets themselves must be relevant, sufficiently representative, and — to the best extent possible — free of errors and complete in view of the intended purpose, with appropriate statistical properties including as regards the persons or groups on whom the system will be used. And they must reflect, to the extent required by the intended purpose, the specific geographical, contextual, behavioural or functional setting of use.
Art 10(5) resolves a real dilemma you will meet repeatedly: to prove your hiring model does not discriminate by ethnicity, you need ethnicity data — which GDPR Art 9 presumptively forbids you to process. The AI Act threads the needle with a narrow, safeguard-wrapped legal gateway: special-category data may be processed for bias detection and correction, and essentially nothing else. It is a compliance tool, not a data-collection licence — every safeguard on the list is a checkable control.
One scope note (Art 10(6)): for high-risk systems not developed through model training, the data-quality duties apply only to the testing datasets — a rules-based Annex III system still has to prove itself on representative test data.
Key terms: data governance, representativeness, special category data, feedback loop, annotation
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.