Acquisition and Selection of Data
Write down where each dataset feeding your AI came from and the criteria you used to decide it was the right data to use.
Plain language
When you build or train an AI system you have to feed it data, and this control says you must record two things for every dataset: where it came from (acquisition) and why you chose it over alternatives (selection). For example, if a recruitment tool is trained on ten years of past hiring records, you note that source, who collected it, and the reasoning that made it a sensible fit, so anyone can later check the data was suitable and you were allowed to use it.
Framework
ISO/IEC 42001:2023
Control effect
Preventative
Classifications
N/A
Official last update
01 Dec 2023
Control Stack last updated
19 June 2026
Official control statement
The organisation shall determine and document details about the acquisition and selection of the data used in AI systems.
Why it matters
If you cannot show where your training data came from, you may discover too late that it was scraped or repurposed without a lawful basis, forcing you to retrain or withdraw the AI system because the model is built on data you had no right to use. A complainant can take a privacy concern to the OAIC, and without acquisition records you cannot demonstrate the data was collected fairly or with consent. Poorly chosen data also bakes in problems you cannot trace: a credit or hiring model trained on an unrepresentative dataset can quietly disadvantage a whole cohort of applicants.
Operational notes
Record acquisition and selection details at the moment a dataset is brought in, before it reaches a training pipeline, because reconstructing provenance (the documented history of where data came from and how it was handled) months later is usually impossible. Keep the record tied to a specific dataset version or snapshot, so when you refresh or re-source data you capture the new origin and selection rationale rather than overwriting the old one. When a dataset is rejected during selection, note why, so the same unsuitable source is not picked up again later.
Implementation tips
- The data lead should keep a dataset register that, for every dataset feeding an AI system, records the origin, who collected it, the date range, and a link to the lawful basis or licence; a simple table per dataset version is enough as long as it is filled in when the data is acquired.
- For data bought or licensed from an outside source, whoever arranges the acquisition should capture the supplier's name, what the data is, and the licence terms that permit AI training, and file that alongside the dataset so the right to use it is documented at the point of acquisition.
- Whoever selects a dataset for a model should write a short selection rationale noting why this data fits the AI system's intended use, for example its coverage, recency, or representativeness, so the choice is traceable rather than implicit.
- The data lead should also log datasets that were considered and rejected, with the reason, so unsuitable sources are not quietly re-introduced later and so selection decisions are documented, not just the data that made the cut.
- Tie each acquisition and selection record to a specific dataset version or snapshot identifier, so that when data is refreshed or re-sourced the new origin and rationale are captured against the new version instead of overwriting the original.
Audit / evidence tips
- AskAsk the data lead for the provenance record of a specific dataset used to train one of the organisation's AI systems.GoodEach dataset has a provenance record naming its source, collector, and time period, and the auditor can trace it to a real origin.
- AskAsk for the documented lawful basis or licence covering the dataset used in that AI system.GoodThe record cites a specific lawful basis or licence that clearly permits using the dataset to train the AI system.
- AskAsk for the selection rationale showing why this dataset was chosen for the AI system.GoodThe selection rationale explains why the dataset fits the intended use and what was considered before choosing it.
- AskAsk to see the dataset register and pick one entry at random.GoodThe register links each dataset version to its documented origin and the reason it was selected.
- AskAsk whether any candidate datasets were considered and then rejected, and for the record of that decision.GoodThere is a record of rejected datasets with clear reasons, showing selection decisions are documented, not just acquisitions.
Cross-framework mappings
How Annex A 7.3 relates to controls across ISO/IEC 27001, ISO/IEC 42001, Essential Eight, and ASD ISM.
ISO 27001
| Control | Notes | Details |
|---|---|---|
handshakeSupports(3)expand_less | ||
| Annex A 5.12 | Annex A 7.3 requires the organisation to document data acquisition and selection for AI systems | |
| Annex A 5.13 | Annex A 7.3 mandates documenting AI data acquisition and selection | |
| Annex A 5.19 | Annex A 7.3 requires documenting how data for AI is acquired and selected | |
These mappings show relationships between controls across frameworks. They do not imply full equivalence or certification.
Related ISO 42001 controls in A.7 Data for AI systems
See all A.7 Data for AI systems controls, or browse the full ISO 42001 Annex A library.
Want to implement this AI control?
Mindset Cyber runs PECB-accredited ISO/IEC 42001 training that maps directly to the AI controls in this library.