Aggregation Bias
Aggregation bias occurs when a trend observed in grouped or combined data is wrongly assumed to hold true for the individual people or subgroups that make up that data. In practice, a pattern that appears in the aggregate may not describe, or may even reverse for, specific individuals or smaller groups. This can lead to inaccurate conclusions when broad findings are applied to particular cases.
Aggregation bias is the systematic inaccuracy in statistical inference that arises when patterns present in aggregated data are incorrectly assumed to apply at the level of the individual units or subgroups from which the data was composed. As commonly defined, it can be induced by the process of grouping data or of combining variables into composite models, and it is a recognized concern in regression analysis where relationships estimated on aggregated data may not reflect relationships at the disaggregated level. Practitioners typically distinguish it from other forms of bias in that its source lies specifically in the aggregation or grouping process rather than in sampling or measurement alone; note that the evidence here addresses the statistical/econometric framing and does not establish a single authoritative definition across all AI or fairness contexts.
Why it matters
Aggregation bias matters because decisions made about individuals or subgroups are frequently informed by patterns observed at a broader, aggregated level. When a relationship holds in combined data but does not hold—or reverses—for the specific units that compose that data, conclusions drawn for particular cases can be systematically inaccurate. In an AI governance and model risk context, this is significant because models are often trained and evaluated on pooled data, and their outputs may then be applied to individuals whose behavior or characteristics diverge from the aggregate trend.
For model risk managers, aggregation bias is a validation concern: a model that performs well against aggregate metrics may nonetheless produce unreliable inferences for subgroups, which can affect both model performance and, depending on context, fairness-related outcomes. Because the source of this bias lies specifically in the grouping or combining of data rather than in sampling or measurement alone, it may not be surfaced by review processes that focus only on data collection or on overall accuracy. Distinguishing it from other bias sources is important for correctly targeting mitigation.
The evidence here addresses the statistical and econometric framing of aggregation bias and does not establish a single authoritative definition across all AI or fairness contexts. Readers applying the concept in a fairness setting should treat the connection to fairness outcomes as context-dependent rather than assumed, and should not conflate the presence of aggregation bias with a determination that a model is unfair or non-compliant under any particular framework.
Who it's relevant to
Inside Aggregation Bias
Common questions
Answers to the questions practitioners most commonly ask about Aggregation Bias.