Valid and Reliable
"Valid" means a test or measurement actually measures what it is supposed to measure, while "reliable" means it produces consistent results when repeated under similar conditions. These are two distinct properties: something can be reliable (consistent) without being valid (correct). A measure that is consistently wrong is reliable but not valid.
As commonly defined in measurement and evaluation literature, validity refers to the degree to which an instrument measures the construct it is intended to measure, whereas reliability refers to the consistency or reproducibility of measurement results across repeated administrations under comparable conditions. The two are related but not interchangeable: reliability is typically treated as a necessary but not sufficient condition for validity, since a measurement may yield reproducible results that are nonetheless not accurate representations of the target construct. A valid measurement is generally reliable, but a reliable measurement is not necessarily valid. Note that these terms originate in research and testing methodology; their application to AI system properties may carry framework-specific meanings that differ from the classical psychometric definitions described here, and that scope is outside this evidence.
Why it matters
The distinction between validity and reliability underpins whether any claim about an AI system's measured performance can be trusted. A metric that is reliable but not valid produces consistent numbers that do not actually reflect the property being assessed, which can create false confidence. As commonly defined in measurement and evaluation literature, reliability is typically treated as a necessary but not sufficient condition for validity, so a measure can be reproducible time and time again while still failing to capture the construct it is meant to represent.
This matters for anyone relying on evaluation results to make governance or risk decisions. If an evaluation instrument consistently returns the same outcome, that consistency alone does not establish that the outcome is correct; a measurement can be reliable and still be systematically wrong. Treating reliability as evidence of validity is a common error, and conflating the two can lead teams to accept measurements that are precise but not accurate representations of the target.
Because these terms originate in research and testing methodology, their use in describing AI system properties may carry framework-specific meanings that differ from the classical definitions described here. Professionals should be careful not to assume that a term labeled "valid and reliable" in one framework maps directly onto the psychometric definitions, and should confirm how a given framework scopes each term before drawing conclusions.
Who it's relevant to
Inside Valid and Reliable
Common questions
Answers to the questions practitioners most commonly ask about Valid and Reliable.