Skip to main content
Category: Risk Assessment & Analysis

Measure Function

Also known as: Measure, Set Function (measure-theoretic)
Simply put

A measure function is a mathematical rule that assigns a number to a set to describe its 'size'—generalizing familiar notions like length, area, and volume. Different measure functions represent different ways of quantifying how big a set is. Note that this is a foundational mathematical concept and is distinct from the AI governance sense of 'measure' (for example, a control or metric used to manage model risk).

Formal definition

In measure theory, a measure is a function that assigns a nonnegative real value (or, in some generalizations, complex values) to sets drawn from a specified collection of subsets, such as a sigma-algebra or delta-ring, subject to axioms including assigning zero to the empty set and, typically, countable additivity over disjoint sets. As commonly defined (for example, in Wolfram MathWorld), a measure m is a nonnegative real function on a delta-ring F satisfying m(emptyset)=0. Each distinct measure embodies a different way to quantify the size of sets, generalizing geometric notions of length, area, and volume. This entry addresses only the mathematical concept and does not cover any AI governance, risk-management, or NIST AI RMF usage of the word 'measure,' which is a separate and unrelated meaning.

Why it matters

The measure function is one of the foundational objects of modern mathematical analysis. By formalizing the intuitive notion of the "size" of a set—generalizing length, area, and volume—it provides the rigorous underpinning for integration theory and, through that, for probability theory and statistics. Much of the quantitative machinery used in data science and machine learning ultimately rests on measure-theoretic foundations, even when practitioners never invoke the term directly.

For readers in AI governance and model risk management, the primary reason this entry matters is disambiguation. The word "measure" appears prominently in risk-management vocabulary—for example, as a control, metric, or activity used to manage model risk—and this usage is entirely separate from the mathematical measure function described here. Confusing the two can lead to imprecise communication between quantitative and governance teams. This entry documents only the mathematical concept so that the distinction remains clear.

Because the value of a measure function depends on which measure is chosen, the concept underscores a broader principle: different rules for quantifying the "size" of sets yield different results. Recognizing that each distinct measure embodies a different way to assess how big a set is helps clarify why measure-theoretic choices are foundational rather than merely technical bookkeeping.

Who it's relevant to

Data scientists and quantitative modelers
Because probability and integration rest on measure-theoretic foundations, this concept underpins much of the mathematics behind statistical and machine learning methods. Practitioners generally encounter it as background theory rather than as a day-to-day tool, but familiarity helps clarify why quantitative choices about the "size" of sets can be foundational.
Model risk and governance professionals
The main relevance here is terminological. The word "measure" carries a distinct meaning in AI governance and risk-management contexts—such as a control or metric used to manage model risk—that is separate and unrelated to the mathematical measure function. This entry addresses only the mathematical sense to prevent the two from being conflated.
Students and readers of mathematical statistics
For those studying measure theory, this entry gives the functional definition: a rule assigning a nonnegative value to sets, subject to axioms including assigning zero to the empty set. It also notes that different measures represent different ways of quantifying set size, a point that is easy to overlook when first encountering the subject.

Inside Measure Function

Measurement Scope Definition
As commonly framed in the NIST AI Risk Management Framework (issued by the U.S. National Institute of Standards and Technology, a voluntary framework), the Measure function centers on identifying which AI risks, impacts, and trustworthiness characteristics will be assessed, and selecting the metrics and methods appropriate to those characteristics.
Quantitative and Qualitative Assessment Methods
The Measure function typically encompasses both numeric measurement (such as performance and error metrics) and qualitative evaluation (such as expert review or structured judgment) applied to identified AI system characteristics. It does not presuppose that all risks are reducible to a single quantitative score.
Trustworthiness Characteristic Evaluation
In many descriptions of the framework, Measure addresses assessment of characteristics such as validity, reliability, safety, security, fairness-related concerns, and explainability. These are evaluated as distinct dimensions rather than collapsed into one aggregate rating.
Testing, Evaluation, Verification, and Validation Activities
Measure commonly draws on TEVV-type activities. Note the expert distinction: verification typically asks whether the system was built correctly against specifications, while validation typically asks whether the system meets its intended use and requirements; the two are related but not interchangeable.
Tracking and Monitoring Inputs
The Measure function generates information about system behavior over time, which supports monitoring for issues such as model performance degradation. Performance degradation is a distinct concept from model risk itself, and Measure supplies observations that feed later management decisions rather than resolving them.
Relationship to Governance and Other Functions
Measure is one function within the broader framework alongside functions oriented to organizational governance, context mapping, and risk response. It operationalizes assessment but does not, by itself, constitute the organizational accountability structures associated with AI governance.

Common questions

Answers to the questions practitioners most commonly ask about Measure Function.

Does the Measure function determine whether an AI system is acceptable to deploy?
Not on its own. The Measure function is commonly framed as the analysis, assessment, benchmarking, and monitoring of AI risks and their impacts—it produces the information and metrics that inform decisions. The judgment about acceptability, tolerance, and go/no-go is typically the province of governance and the Manage function, which act on Measure's outputs. Treating Measure as the decision point rather than the measurement point is a frequent conflation.
Is measuring a model's performance the same as measuring its risk under the Measure function?
No, and professionals should be careful here. Performance metrics (such as accuracy or error rates) are one input, but the Measure function as commonly described is broader: it addresses trustworthiness characteristics, potential impacts, and risks that may not be captured by performance alone. A model can perform well on its stated metric while still presenting risks the Measure function is intended to surface. Equating the two understates what measurement is meant to cover.
What kinds of methods and metrics are typically used within the Measure function?
Approaches commonly include quantitative metrics, qualitative assessments, benchmarking against defined baselines, and ongoing monitoring. Selection typically depends on the system's context, the risks previously identified, and the trustworthiness characteristics being evaluated. There is no single mandated metric set; the appropriate methods vary by use case, and organizations generally document why chosen methods are fit for purpose.
How does the Measure function connect to the risks identified earlier in a risk process?
In many framings, measurement operationalizes the risks and impacts previously identified—turning them into things that can be assessed and tracked. Risks that were catalogued but never assigned measurement methods can go unmonitored, so practitioners often trace each identified risk to a corresponding measurement approach and note where a risk cannot be reliably measured.
How often should measurement activities be repeated after deployment?
Measurement is commonly treated as ongoing rather than a one-time exercise, given that model behavior, data, and operating conditions can change over time. Cadence typically depends on the system's risk level, its rate of change, and organizational policy. Documenting the chosen frequency and the triggers for re-measurement is a common practice; this entry does not specify required intervals, as these are context-dependent.
What should be done when a risk is difficult or impossible to measure reliably?
Where reliable measurement is not feasible, practitioners often document the limitation explicitly, note the residual uncertainty, and escalate it to governance and the Manage function so that the gap is accounted for in decision-making. Recognizing and recording the boundaries of what can be measured is generally treated as part of sound practice rather than a failure to be hidden.

Common misconceptions

The Measure function eliminates or fully quantifies AI risk once metrics are in place.
Measurement is a means of assessing and characterizing risk, not eliminating it. As commonly framed, the Measure function reduces uncertainty and informs risk response, but residual risk typically remains and some risks may resist reliable quantification.
The Measure function is the same as model validation under banking model risk guidance such as SR 11-7 (issued by the U.S. Federal Reserve and OCC).
The NIST AI RMF and SR 11-7 originate from different bodies, apply in different contexts, and are not interchangeable; SR 11-7 is supervisory guidance oriented to banking model risk management, while the NIST AI RMF is a voluntary framework. They may overlap in assessment concepts, but the Measure function should not be equated with a specific regulatory validation requirement.
Measuring a system's performance is equivalent to measuring its fairness or its overall risk.
Model performance, bias, fairness, and overall model risk are distinct concepts. High measured performance does not establish fairness, and bias metrics do not by themselves settle contested fairness questions, which often depend on context and normative choices.

Best practices

Select measurement methods and metrics that map explicitly to the specific trustworthiness characteristics and risks identified for the system, rather than defaulting to a single aggregate performance number.
Combine quantitative metrics with qualitative and expert assessment, and document the assumptions and limitations of each method so downstream users understand what the measurements do and do not establish.
Distinguish clearly in documentation between verification activities and validation activities, and between performance measurement and risk assessment, to avoid conflating related concepts.
Establish ongoing monitoring so that measurement is repeated over the system lifecycle to detect performance degradation and changing conditions, rather than treating measurement as a one-time exercise.
Record contested or context-dependent measurement choices, such as fairness definitions or thresholds, explicitly so that reviewers and second-line functions can evaluate the basis for those decisions.
Feed measurement outputs into the organization's risk response and governance processes, recognizing that measurement informs but does not by itself manage or eliminate residual risk.