Data for Development and Enhancement of AI System
Set up, write down and actually follow clear processes for managing the data you use to build and improve your AI systems.
Plain language
When you build or improve an artificial intelligence (AI) system, it learns from data. This control says you must have proper, written-down processes for how that development data is handled at every step: how it is collected, where it is stored, how it is labelled or transformed, how versions are tracked, and how long it is kept. "Define, document and implement" means three things: agree the steps, write them down so they are not just in someone's head, and then genuinely work to them. The reason this matters is that the data is what shapes the AI. If different people on a project collect and prepare training data in their own ad hoc ways, you end up with an AI whose behaviour you cannot explain or rebuild. A worked example: a team trains a pricing model on a dataset, the model goes live, and months later the model starts behaving oddly. With documented data processes you can pull the exact dataset version, see how it was assembled, and fix the problem. Without them, nobody can reconstruct what the model was trained on, so you are stuck. This control is about the development data pipeline itself, not about teaching staff or running risk workshops.
Framework
ISO/IEC 42001:2023
Control effect
Preventative
Classifications
N/A
Official last update
01 Dec 2023
Control Stack last updated
19 June 2026
Official control statement
The organisation shall define, document and implement data management processes related to the development of AI systems.
Why it matters
Without defined and documented data-management processes, the dataset that trained a model often cannot be reconstructed: nobody recorded which records were pulled, how they were cleaned, or which version went into the model. When that model later misbehaves or a regulator asks you to justify it, you cannot retrain it on the same basis, trace its data lineage (the documented chain showing where each piece of data came from and how it was transformed), or even prove what it learned from. The system effectively has to be rebuilt from scratch or pulled from use.
Operational notes
Treat the documented data process as a living pipeline, not a one-off write-up: version each development dataset so you can always point to the exact data a given model was trained on. Whenever a new data source is added, a collection or labelling method changes, or a model is retrained, update the process and the dataset records together rather than waiting for an annual review. Keep older dataset versions and their documentation long enough to reproduce or investigate any model still in use.
Implementation tips
- The data or AI lead should write a single data-management process for AI development that spells out, step by step, how data is collected, stored, labelled or transformed, versioned and retained, so every project handles development data the same way instead of improvising.
- Whoever assembles training data should version each development dataset and record which version feeds each model, for example by tagging the dataset and noting that tag against the trained model, so you can always point back to the exact data a model learned from.
- The data lead should keep lineage records that capture where each dataset came from and the cleaning, filtering and labelling steps applied, so a training dataset can be reproduced rather than reconstructed from guesswork later.
- The data or AI lead should make the documented process easy to follow in practice by giving teams a short intake or preparation checklist to complete for each dataset, turning the written process into something that is actually implemented, not just filed.
- Whoever owns the process should update both the process and the dataset documentation whenever a new data source is brought in, a collection or labelling method changes, or a model is retrained, so the records keep matching what the development pipeline really does.
Audit / evidence tips
- AskAsk for the documented data-management process used when developing or enhancing AI systems.GoodThere is a written process that sets out concrete steps for handling development data from collection through to retention, with an owner and a last-reviewed date.
- AskAsk for the dataset register or version log and pick one live AI model from it.GoodEach AI model can be tied to a specific, dated dataset version that the organisation can still produce.
- AskAsk to see the data lineage or transformation records for one development dataset.GoodThe transformation history is recorded clearly enough that someone could rebuild the same training dataset from the source data.
- AskAsk the AI or data lead to walk through how the documented process was followed on a recent development project.GoodThe named process was demonstrably applied on a recent project, with completed records matching its steps.
- AskAsk for the change history of the data process and dataset documentation.GoodThe process and dataset records show dated updates triggered by real data or model changes.
Cross-framework mappings
How Annex A 7.2 relates to controls across ISO/IEC 27001, ISO/IEC 42001, Essential Eight, and ASD ISM.
ISO 27001
| Control | Notes | Details |
|---|---|---|
handshakeSupports(3)expand_less | ||
| Annex A 5.10 | Annex A 7.2 requires the organisation to implement defined processes for managing data used in AI development and enhancement | |
| Annex A 5.12 | Annex A 7.2 requires data management processes for AI system development and enhancement, including governance over what data is used and... | |
| Annex A 5.37 | Annex A 7.2 requires the organisation to define, document and implement data management processes for developing and enhancing AI systems... | |
These mappings show relationships between controls across frameworks. They do not imply full equivalence or certification.
Related ISO 42001 controls in A.7 Data for AI systems
See all A.7 Data for AI systems controls, or browse the full ISO 42001 Annex A library.
Want to implement this AI control?
Mindset Cyber runs PECB-accredited ISO/IEC 42001 training that maps directly to the AI controls in this library.