Skip to main content
Category: Data Governance & Quality

Data Governance

Simply put

Data governance is the set of rules, roles, and processes an organization uses to manage its data throughout its life, from when data is collected to when it is securely disposed of. Its aim is to help ensure that data is reliable, consistent, secure, and can be trusted for use. It typically combines policies, defined responsibilities, and supporting technology tools.

Formal definition

As commonly defined across vendor and practitioner sources, data governance is a principled, life-cycle-oriented framework comprising policies, procedures, standards, roles, metrics, and technology tools for managing an organization's data assets from acquisition and ingestion through use, analytics, and secure disposal. Its stated objectives generally include ensuring that data is reliable, consistent, secure, trustworthy, and used effectively and efficiently. Note that the evidence provided consists largely of vendor and community descriptions rather than regulatory or standards-body definitions; specific control requirements, maturity models, and role structures vary by source, sector, and jurisdiction, and are out of scope here. Data governance should be distinguished from AI governance (organizational oversight of AI systems) and from model risk management, though it frequently overlaps with both as a data-quality and accountability foundation.

Why it matters

Data governance matters because the reliability of nearly every downstream analytical, operational, and AI-driven decision depends on the quality and trustworthiness of the underlying data. As commonly described across practitioner and vendor sources, governance provides the policies, roles, and processes that help ensure data is consistent, secure, and can be trusted throughout its life cycle. Without such structures, organizations risk making decisions on data that is inconsistent, poorly controlled, or of unknown provenance, which can undermine both operational outcomes and accountability.

For practitioners in AI governance and model risk management, data governance functions as a foundational layer rather than a substitute for either discipline. Reliable data supports, but does not by itself constitute, sound model development, validation, or organizational oversight of AI systems. It is important to keep these distinctions in view: data governance addresses the management of data assets, whereas AI governance addresses organizational oversight of AI systems, and model risk management addresses the identification, measurement, monitoring, and control of risks arising from model use. The three frequently overlap because data quality and accountability underpin trustworthy models, but they should not be collapsed into one another.

A note on scope and limitations: the descriptions summarized here draw largely from vendor and community sources rather than regulatory or standards-body definitions. Specific control requirements, maturity models, and role structures vary by source, sector, and jurisdiction. Readers should treat data governance as a broadly recognized practice with meaningful implementation differences, not as a single standardized framework with universally required controls.

Who it's relevant to

Data Scientists and Model Developers
Data governance supplies the reliability, consistency, and provenance controls that model development typically depends on. Practitioners should treat well-governed data as a supporting foundation for model work rather than as a guarantee of model quality, since governance of data assets is distinct from validation of the models built on that data.
Model Risk Managers
Data quality and accountability are commonly regarded as inputs to sound model risk management, but data governance and model risk management remain distinct disciplines. Model risk managers may rely on data governance as a control foundation while separately addressing the identification, measurement, monitoring, and control of risks arising from model use.
AI Governance and Policy Specialists
Because data governance frequently underpins trustworthy AI systems, those responsible for organizational oversight of AI should understand where the two intersect. Data governance addresses management of data assets, while AI governance addresses oversight of AI systems; the overlap is meaningful but the two should not be conflated.
Auditors and Compliance Officers
Data governance defines policies, roles, and processes that can be examined for evidence of accountability and control over data assets. Auditors should note that specific control requirements and role structures vary by source, sector, and jurisdiction, and that the descriptions available here derive largely from vendor and practitioner sources rather than binding regulatory definitions.

Inside Data Governance

Data Quality Management
Processes and controls intended to maintain the accuracy, completeness, consistency, timeliness, and validity of data used across the organization, including for model development and monitoring. In model risk contexts, data quality issues in training or input data are a recognized source of model risk.
Data Ownership and Stewardship
The assignment of accountability for defined data assets, typically distinguishing data owners (accountable for a domain of data) from data stewards (responsible for day-to-day quality and definition maintenance). This supports the accountability structures central to AI governance without, by itself, constituting model risk management.
Data Lineage and Provenance
Documentation of where data originates, how it flows through systems, and how it is transformed over time. Lineage supports traceability, reproducibility, and the ability to assess whether data is fit for a given use, and it is often cited as an input to model validation and audit.
Policies, Standards, and Procedures
The written rules governing how data is defined, classified, accessed, retained, and disposed of. These typically sit within the broader governance layer that assigns oversight and decision rights, as distinct from the risk measurement and control activities of model risk management.
Access Control and Privacy Management
Controls governing who may view or use data and under what conditions, often aligned with applicable privacy obligations. The specific legal requirements vary by jurisdiction and sector, so scope should be assessed against the frameworks actually applicable to the organization.
Metadata Management and Data Cataloguing
The capture and maintenance of descriptive information about data assets (definitions, formats, classifications, business meaning), commonly maintained in a catalogue to make data discoverable and to support consistent interpretation.
Data Classification and Retention
The categorization of data by sensitivity or criticality and the rules governing how long it is kept and when it is disposed of, supporting both compliance and operational control.

Common questions

Answers to the questions practitioners most commonly ask about Data Governance.

Is data governance the same thing as data management?
No, though the two are frequently conflated. Data governance typically refers to the organizational structures, policies, roles, and accountability that determine how data is treated as an asset—who decides, who is responsible, and against what standards. Data management refers to the operational activities that execute those decisions, such as storage, integration, and processing. As commonly defined, governance sets the rules and oversight while management carries them out; treating them as interchangeable tends to obscure where accountability actually sits.
Does data governance overlap with AI governance and model risk management, or are they separate?
They overlap without being the same. AI governance concerns the organizational oversight and accountability for AI systems as a whole, while model risk management focuses on identifying, measuring, monitoring, and controlling the risks arising from model use. Data governance is a related discipline addressing the quality, lineage, access, and stewardship of data, which often feeds into both. Data issues frequently surface as inputs to model risk (for example, through data quality affecting model performance), but data governance is a distinct function with its own scope and should not be collapsed into either AI governance or model risk management.
Who typically owns data governance responsibilities within an organization?
Ownership is commonly distributed rather than held by a single role. Many organizations assign accountability to data owners or stewards for specific data domains, while a governance council or committee sets enterprise-wide policy. In frameworks that use a three-lines-of-defense model, business units and data stewards often act in the first line, a governance or risk oversight function in the second, and audit in the third. The precise allocation varies by organization and sector, so titles and responsibilities should be confirmed against your own operating model rather than assumed.
How does data governance support model risk management in practice?
In many programs, data governance provides the documented data lineage, quality standards, and access controls that model risk activities rely on. For example, validation work typically examines whether input data is fit for purpose, and governance artifacts such as data dictionaries and quality metrics can support that examination. It is important to note that strong data governance reduces and helps manage data-related risk to models but does not eliminate it; the specific expectations and integration points depend on your regulatory environment and internal policy.
What controls are commonly used to operationalize data governance?
Commonly cited controls include data quality rules and monitoring, access and permissioning controls, documented data lineage, metadata and cataloging practices, retention and disposal policies, and defined stewardship roles. The appropriate mix depends on the organization's risk profile, data sensitivity, and applicable regulatory obligations. This list is illustrative rather than exhaustive, and the presence of controls does not by itself demonstrate their effectiveness—monitoring and review are typically needed to confirm they operate as intended.
How can an organization measure whether its data governance is effective?
Effectiveness is typically assessed through a combination of indicators rather than a single metric—examples include data quality measures, remediation timeliness, clarity of ownership, and evidence that policies are followed in practice. Because there is no universally authoritative set of metrics, organizations often tailor indicators to their objectives and regulatory context. Independent review, such as internal audit in a third-line role, is commonly used to test whether governance is functioning as designed rather than existing only on paper.

Common misconceptions

Data governance and model risk management are the same discipline.
Data governance is primarily concerned with the organizational structures, policies, ownership, and oversight that ensure data is managed as a controlled asset. Model risk management, as historically framed by guidance such as SR 11-7 (issued in the U.S. by the Federal Reserve and OCC), focuses on identifying, measuring, monitoring, and controlling risks arising from model use. They overlap, because poor data quality is a recognized contributor to model risk, but they remain distinct disciplines and should not be collapsed into one.
Strong data governance guarantees high-quality models or eliminates data-related risk.
Data governance measures reduce and help manage data-related risk but do not eliminate it. Well-governed data can still be unrepresentative, biased, or unsuitable for a particular modeling purpose, and governance controls do not substitute for validation and ongoing monitoring of how data performs in a specific model.
Data governance requirements are uniform across jurisdictions and sectors.
The specific obligations tied to data, particularly around privacy, access, and retention, vary by jurisdiction and by sector, and the treatment of data within broader AI-related frameworks continues to evolve. Requirements should be scoped to the frameworks actually applicable to the organization rather than assumed to be universal.

Best practices

Assign clear data ownership and stewardship roles so that accountability for each defined data domain is explicit, and align these roles with the organization's broader governance and lines-of-defense structure.
Maintain data lineage and provenance documentation for data used in models so that inputs can be traced, reproduced, and assessed for fitness during validation and audit.
Establish and enforce data quality controls (covering accuracy, completeness, consistency, timeliness, and validity) and treat identified data quality issues as a potential source of model risk that is escalated accordingly.
Maintain a metadata catalogue with agreed definitions and classifications so that data meaning is interpreted consistently across teams and use cases.
Scope access control, privacy, and retention policies to the legal and regulatory frameworks actually applicable to the organization and its sector, rather than applying a single assumed standard universally.
Coordinate data governance activities with model risk management functions so that data quality and lineage findings feed into model validation and ongoing monitoring, while keeping the two functions' responsibilities distinct.