Data Lineage
Data lineage is a record and visual map of a dataset's journey, showing where the data came from, how it changed as it moved through systems, and where it ends up being used. It helps organizations understand and trace the history of their data over time. Think of it as a documented trail that follows data from its original source to its final point of consumption.
Data lineage is the process of recording, tracking, and visualizing the flow and transformation of data (and, in some framings, AI artifacts) across its life cycle, from origin through intermediate transformations to consumption. As commonly described by data platform vendors, it captures a dataset's sources, the transformations applied as it moves through pipelines and systems, and its current location or end use, typically rendered as a map or graph of dependencies. Note that the evidence provided reflects vendor and industry descriptions rather than a single authoritative or regulatory definition; specific implementations, granularity (e.g., column-level versus table-level), and scope vary by tool and organizational context. This entry does not address how data lineage maps to particular regulatory recordkeeping obligations, which fall outside the supplied evidence.
Why it matters
Data lineage matters because organizations cannot govern, validate, or trust data they cannot trace. When a dataset feeds an AI model, a report, or an operational decision, understanding where that data originated, how it was transformed, and where it is now consumed is foundational to assessing its quality and fitness for purpose. Without a documented trail, errors introduced early in a pipeline can propagate silently, and teams may be unable to determine whether a downstream problem stems from the source data, an intermediate transformation, or the point of use.
In the context of AI governance and model risk management, lineage supports several distinct activities that professionals should not conflate. It can help establish the provenance of training and input data, contribute to reproducibility, and inform impact analysis when a source system changes. That said, data lineage is a supporting capability rather than a control that eliminates risk on its own; it makes data flows visible and traceable, which in turn enables review, but it does not by itself validate that data is accurate or that a model is fit for use.
The evidence supplied here reflects vendor and industry descriptions rather than a single authoritative or regulatory definition, and the specific granularity and scope of lineage vary by tool and organizational context. Readers should therefore treat lineage as a broadly recognized data governance practice whose exact implementation and its mapping to particular recordkeeping or regulatory obligations depend on context not covered by this evidence.
Who it's relevant to
Inside Data Lineage
Common questions
Answers to the questions practitioners most commonly ask about Data Lineage.