Skip to content
arrow_back
Annex A 7.3psychologyISO/IEC 42001:2023

Acquisition and Selection of Data

Write down where each dataset feeding your AI came from and the criteria you used to decide it was the right data to use.

record_voice_over

Plain language

When you build or train an AI system you have to feed it data, and this control says you must record two things for every dataset: where it came from (acquisition) and why you chose it over alternatives (selection). For example, if a recruitment tool is trained on ten years of past hiring records, you note that source, who collected it, and the reasoning that made it a sensible fit, so anyone can later check the data was suitable and you were allowed to use it.

Framework

ISO/IEC 42001:2023

Control effect

Preventative

Classifications

N/A

Official last update

01 Dec 2023

Control Stack last updated

19 June 2026

Official control statement

The organisation shall determine and document details about the acquisition and selection of the data used in AI systems.
psychologyISO/IEC 42001:2023Annex A 7.3
priority_high

Why it matters

If you cannot show where your training data came from, you may discover too late that it was scraped or repurposed without a lawful basis, forcing you to retrain or withdraw the AI system because the model is built on data you had no right to use. A complainant can take a privacy concern to the OAIC, and without acquisition records you cannot demonstrate the data was collected fairly or with consent. Poorly chosen data also bakes in problems you cannot trace: a credit or hiring model trained on an unrepresentative dataset can quietly disadvantage a whole cohort of applicants.

settings

Operational notes

Record acquisition and selection details at the moment a dataset is brought in, before it reaches a training pipeline, because reconstructing provenance (the documented history of where data came from and how it was handled) months later is usually impossible. Keep the record tied to a specific dataset version or snapshot, so when you refresh or re-source data you capture the new origin and selection rationale rather than overwriting the old one. When a dataset is rejected during selection, note why, so the same unsuitable source is not picked up again later.

build

Implementation tips

  • The data lead should keep a dataset register that, for every dataset feeding an AI system, records the origin, who collected it, the date range, and a link to the lawful basis or licence; a simple table per dataset version is enough as long as it is filled in when the data is acquired.
  • For data bought or licensed from an outside source, whoever arranges the acquisition should capture the supplier's name, what the data is, and the licence terms that permit AI training, and file that alongside the dataset so the right to use it is documented at the point of acquisition.
  • Whoever selects a dataset for a model should write a short selection rationale noting why this data fits the AI system's intended use, for example its coverage, recency, or representativeness, so the choice is traceable rather than implicit.
  • The data lead should also log datasets that were considered and rejected, with the reason, so unsuitable sources are not quietly re-introduced later and so selection decisions are documented, not just the data that made the cut.
  • Tie each acquisition and selection record to a specific dataset version or snapshot identifier, so that when data is refreshed or re-sourced the new origin and rationale are captured against the new version instead of overwriting the original.
fact_check

Audit / evidence tips

  • AskAsk the data lead for the provenance record of a specific dataset used to train one of the organisation's AI systems.GoodEach dataset has a provenance record naming its source, collector, and time period, and the auditor can trace it to a real origin.
  • AskAsk for the documented lawful basis or licence covering the dataset used in that AI system.GoodThe record cites a specific lawful basis or licence that clearly permits using the dataset to train the AI system.
  • AskAsk for the selection rationale showing why this dataset was chosen for the AI system.GoodThe selection rationale explains why the dataset fits the intended use and what was considered before choosing it.
  • AskAsk to see the dataset register and pick one entry at random.GoodThe register links each dataset version to its documented origin and the reason it was selected.
  • AskAsk whether any candidate datasets were considered and then rejected, and for the record of that decision.GoodThere is a record of rejected datasets with clear reasons, showing selection decisions are documented, not just acquisitions.
link

Cross-framework mappings

How Annex A 7.3 relates to controls across ISO/IEC 27001, ISO/IEC 42001, Essential Eight, and ASD ISM.

ISO 27001

ControlNotesDetails
handshakeSupports(3)expand_less
Annex A 5.12Annex A 7.3 requires the organisation to document data acquisition and selection for AI systems
Annex A 5.13Annex A 7.3 mandates documenting AI data acquisition and selection
Annex A 5.19Annex A 7.3 requires documenting how data for AI is acquired and selected

These mappings show relationships between controls across frameworks. They do not imply full equivalence or certification.

See all A.7 Data for AI systems controls, or browse the full ISO 42001 Annex A library.

psychology

Want to implement this AI control?

Mindset Cyber runs PECB-accredited ISO/IEC 42001 training that maps directly to the AI controls in this library.

Mapping detail

Mapping

Direction

Controls