THE CORE AIM OF VOCABULARY MASTERY · AI DATA PROVENANCE · SOURCE → COLLECTION → TRANSFORMATION → DATASET → MODEL
Did you know a dataset can look perfectly organised while nobody can answer where its information originally came from? AI data provenance vocabulary gives us the language to investigate that question. Terms such as provenance, origin, source registry, lineage, transformation, version, attribution, licence and audit trail help trace how information enters an AI system and changes along the way.
The core aim of vocabulary mastery for AI data provenance is traceability. Readers should be able to identify where a dataset or record came from, who or what collected it, what processing altered it, which version was used for training or retrieval, what rights apply and whether a later output can be traced back to reliable inputs.
This article is part of the eduKateSG Vocabulary Hub. For complementary learning, see Data Governance Vocabulary, Data Quality Vocabulary and Metadata Vocabulary. Provenance connects these disciplines through evidence of origin and change.
The central idea: Knowing what a data item says is useful. Knowing where it came from, how it changed and whether you are entitled to use it is what makes the record accountable.
The 60-second provenance vocabulary router
- Origin: source, creator, collector, publication date.
- Identity: dataset ID, document ID, version, checksum.
- Process: extraction, cleaning, transformation, annotation.
- Connection: lineage, derived-from, used-by, generated-by.
- Responsibility: owner, agent, steward, reviewer.
- Rights and verification: licence, attribution, consent, audit trail.
Ten useful AI data provenance terms, explained
| Term | Practical meaning | The question it answers |
|---|---|---|
| Provenance | Information about an item’s origin, history and the processes that produced or affected it. | How did this information arrive here? |
| Lineage | A record of relationships among upstream and downstream datasets or outputs. | What was derived from what? |
| Source registry | An organised record of sources and their identifiers, authority and relevant constraints. | Which source is this? |
| Metadata | Data describing other data, such as source date, format, schema or steward. | What do we know about the record? |
| Transformation | An operation that changes data, such as normalisation, filtering or joining. | What changed on the way? |
| Version | An identified state of a document, dataset or model at a particular point. | Which copy was used? |
| Checksum | A digest used to detect whether digital content has changed. | Are these bytes the same? |
| Attribution | Credit identifying a source or creator where relevant. | Who should be acknowledged? |
| Licence | The terms governing permitted use of material. | What reuse is authorised? |
| Audit trail | A documented sequence of actions or changes. | Can the process be reconstructed? |
A worked example: a school science knowledge assistant
Imagine a school creates a science question-answering assistant using curriculum documents, open educational resources and local teaching notes. A student asks why condensation forms on a cold glass. The answer may sound correct, but a trustworthy content workflow also knows which source taught that explanation, whether it is suitable for the student’s level and when the source was last checked.
Now suppose a curriculum document is revised. Without provenance, the team may have no reliable way to identify which passages, embeddings or answers were derived from the older version. With provenance, the change can be followed: source document → extracted text → cleaned passage → retrieval index → generated answer. Each step has an identity and a history.
Provenance, lineage and metadata are not synonyms
Provenance: the history of an item
Provenance asks about origin, derivation and responsibility. It can describe the people, organisations, programs and activities involved in producing or changing an item.
Lineage: the dependency path
Lineage concentrates on relationships between sources and outputs. If a spreadsheet, data warehouse table and search index all originate from the same source, lineage helps identify the downstream products affected when that source changes.
Metadata: the descriptive information
Metadata can include titles, timestamps, creators, identifiers, licences and technical formats. Some metadata supports provenance directly, but a list of file properties alone may not tell the whole derivation story. A well-designed system uses all three concepts together.
The W3C provenance model: entity, activity and agent
The W3C PROV family provides a useful conceptual vocabulary for provenance. An entity is a thing such as a document, dataset or model artifact. An activity is a process such as collecting, editing or converting it. An agent is a person, organisation or software actor associated with an activity or responsibility.
In our school example, a syllabus PDF is an entity. Extracting its text is an activity. The extraction service or responsible organisation can be represented as an agent. The cleaned text becomes another entity derived through that activity. This is a clearer explanation than “the data was processed.”
Provenance and evidence quality are different
A complete chain of custody does not prove that a statement is true. A record can be perfectly traceable to an inaccurate source. Conversely, a true claim with no traceable source may still be difficult to verify. Source authority, factual accuracy, rights, freshness and provenance each answer separate questions.
Provenance is not automatically proof of permission
The presence of a licence name or copyright notice in metadata does not itself authorise every use. A dataset might permit research but restrict commercial reuse or redistribution. Content used in AI pipelines can also be subject to contractual conditions, privacy obligations and other rules. Record the relevant terms, their source and the date reviewed; do not infer permission simply because material is publicly accessible.
Data versioning and reproducibility
Suppose a model was trained on “dataset-final.csv.” Three weeks later the file is edited but retains the same name. Reproducing the training run may become impossible. Version identifiers, checksums, source dates and transformation settings help preserve the actual state used.
This matters equally for retrieval systems. When an answer depends on a document updated yesterday, the system should ideally distinguish the older indexed copy from the newly published document and determine whether re-indexing or review is required.
How to build a simple provenance record
- Identify: give the source a stable name, URI or document identifier.
- Describe: record creator, publisher, source type, access date and version.
- Authorise: note the licence or permission and any restrictions requiring review.
- Transform: record each important extraction, cleaning, filtering or annotation step.
- Link: connect transformed outputs to their upstream sources.
- Verify: use hashes, sample checks and quality controls where appropriate.
- Update: flag changed, withdrawn, superseded or expired sources.
A practical exercise: trace one teaching note
Choose a short public learning resource. Record its URL, title, publisher, publication or revision date and relevant usage conditions. Summarise one paragraph in your own words and note that your summary was derived from the source. Next, imagine the publisher changes a key definition. Which notes, explanations and answers would need reviewing? If you can locate them, you are already practising provenance thinking.
Common AI data provenance mistakes
- “The URL is the provenance.” Repair: record source identity, version, derivation and responsible processes as needed.
- “A checksum verifies accuracy.” Repair: a checksum helps detect changes in bytes; it does not establish factual correctness.
- “Public means free to reuse.” Repair: review licence, privacy and other applicable permissions.
- “All cleaned data is original data.” Repair: preserve links from transformed records to their sources.
- “The latest version always explains the old model.” Repair: retain the actual data and model versions used at training time.
Frequently asked questions
What is AI data provenance?
It is information about the origins, transformations, responsibility and relevant history of data used in or produced by AI systems.
What is the difference between provenance and lineage?
Provenance broadly records how and by whom an item came to exist; lineage specifically tracks its connections to upstream and downstream data.
Is data provenance the same as citing sources?
No. Citations identify evidence for claims; provenance can additionally describe versions, processing activities, dependencies and responsible actors.
Why do AI models need data provenance?
To support reproducibility, data-quality investigation, rights review, correction of outdated information and accountable deployment.
Can a dataset have provenance but still be unreliable?
Yes. Provenance makes a source traceable, not automatically correct or appropriate.
Authoritative reference
The W3C PROV Model Primer introduces entities, activities and agents and explains how provenance records can represent the origin and transformation of digital information.
Where this fits in the eduKateSG ecosystem
- Vocabulary Hub
- Data Governance Vocabulary
- Metadata Vocabulary
- Data Catalog Vocabulary
- Data Quality Vocabulary
The AI data provenance vocabulary standard
Mastery is reached when a learner can trace a result to its specific source and version, explain the transformations that followed, identify relevant responsibility and usage conditions, and describe what must be rechecked when the source changes. That is how information becomes more intelligible, correctable and reusable.