Skip to main content
Category: Risk Classification & Tiering

Model Risk Rating

Also known as: Model Tiering, Model Risk Tiering, Model Risk Scorecard
Simply put

A model risk rating is a label or score that an organization assigns to a model to indicate how much risk it could pose if it performs poorly or is used incorrectly. It helps the organization decide how much scrutiny, testing, and oversight a given model needs, with higher-risk models typically receiving more attention. The term is sometimes also called model tiering.

Formal definition

A model risk rating is a categorization or score assigned to an individual model to express its relative level of model risk, commonly used to prioritize and calibrate the intensity of validation, monitoring, and governance activities. In many frameworks the rating is derived from a set of risk attributes assessed through a qualitative or scorecard-based approach, and the practice is frequently termed model tiering. Rating criteria and scales are typically institution-specific rather than fixed by a single authoritative definition, and this entry describes model risk rating as a risk-prioritization mechanism rather than a measure of model performance itself; the two are distinct, since a rating reflects potential adverse consequences of incorrect or misused models rather than observed accuracy.

Why it matters

Model risk rating matters because organizations rarely have the resources to subject every model to the same depth of validation, monitoring, and oversight. By assigning each model a rating or tier that reflects the potential adverse consequences of incorrect or misused models, an institution can prioritize where to concentrate scrutiny, directing more intensive testing and governance toward models whose failure would cause the greatest harm. This risk-based prioritization is a core mechanism for allocating limited model risk management effort in a defensible, consistent way.

A rating also supports accountability and internal consistency. Because rating criteria and scales are typically institution-specific rather than fixed by a single authoritative definition, a documented rating approach gives an organization a repeatable basis for explaining why a given model received the level of oversight it did. It is important to note that a model risk rating reflects the potential adverse consequences of a model performing poorly or being used incorrectly, not the model's observed accuracy or performance. Treating a rating as a measure of how well a model performs is a common error; the two are distinct, and conflating them can lead an organization to under-scrutinize a poorly rated but high-consequence model.

The practical value of a rating depends on how well the underlying criteria capture the ways a model could cause harm. Ratings are a means of reducing and managing model risk by focusing oversight, not a control that eliminates it; a model rated high risk still requires the actual validation, monitoring, and governance activities that the rating is meant to prioritize.

Who it's relevant to

Model Risk Managers
Model risk managers typically own the rating methodology and use it to calibrate the intensity of validation, monitoring, and governance across a model inventory. They are responsible for ensuring the rating criteria capture potential adverse consequences of incorrect or misused models and are applied consistently.
Model Validators
Validators often use a model's risk rating to determine the depth and frequency of validation activities, concentrating effort on higher-rated models. They should keep in mind that the rating reflects potential consequences rather than observed performance, so validation remains necessary regardless of tier.
Compliance Officers and Auditors
Compliance and audit professionals rely on documented, consistently applied ratings as evidence that oversight is allocated on a defensible, risk-based basis. Because rating scales are usually institution-specific, they typically assess whether the methodology is reasonable and applied as designed rather than against a single universal standard.
Model Owners and Developers
Those who build and deploy models are affected because a model's rating determines the governance obligations, testing, and monitoring it will be subject to. Understanding the attributes that drive a rating helps them anticipate the level of scrutiny a model will face.
Senior Management and Governance Committees
Leaders responsible for AI governance and oversight structures use ratings to understand where the highest-consequence models sit and to prioritize attention and resources. Ratings support, but do not replace, the organizational accountability that governance is intended to provide.

Inside Model Risk Rating

Materiality or Exposure Component
An assessment of the potential magnitude of impact if the model produces erroneous or misused output, often reflecting factors such as financial exposure, volume of decisions, or affected populations. This component addresses how much is at stake rather than how the model was built.
Model Complexity Component
An evaluation of the sophistication and opacity of the model's methodology, which can influence how difficult it is to validate, interpret, or explain. Higher complexity does not automatically mean higher rating, but it is commonly treated as a contributing factor.
Uncertainty or Data Reliability Component
Consideration of the quality, representativeness, and stability of inputs and assumptions, as well as the degree of confidence in the model's outputs. Weaknesses here typically elevate the assessed risk.
Use and Reliance Context
How the model output is used, the degree of human review or automation, and the extent to which decisions depend on the model. The same model may warrant different ratings depending on its use context.
Rating Scale and Categorization
A defined scheme, often tiered (for example, high, medium, low), used to summarize the assessed level of model risk. The specific scales and thresholds are set by each institution and are not standardized across the industry.

Common questions

Answers to the questions practitioners most commonly ask about Model Risk Rating.

Is a model's risk rating the same as its performance score or accuracy metric?
No. A model risk rating is not a measure of how well a model performs. Performance metrics (such as accuracy or error rates) describe model quality, whereas a risk rating typically reflects the potential for adverse consequences arising from model use, incorporating factors such as materiality of the decisions the model supports, complexity, and reliance placed on outputs. A highly accurate model can still carry a high risk rating if it drives material decisions, and a lower-performing model may be rated lower risk if its use is limited. Conflating the two is a common error.
Does assigning a high risk rating and applying controls eliminate the model's risk?
No. A risk rating is a classification used to prioritize oversight and calibrate control intensity; it does not remove risk. Controls applied to higher-rated models are intended to reduce and manage risk, not eliminate it. Even after controls are applied, residual risk typically remains. Professionals should be careful to distinguish inherent risk (before controls) from residual risk (after controls) and avoid implying that a rating or its associated controls make a model risk-free.
What factors are commonly used to assign a model risk rating?
In many frameworks, ratings draw on factors such as the materiality or financial impact of decisions the model informs, model complexity, the degree of reliance on model outputs, data quality and availability, and the potential consequences of model failure. The specific factors and their weighting vary by institution and, in regulated settings such as banking, may be shaped by supervisory expectations. There is no single universally mandated set of factors, so organizations should document their chosen criteria and rationale.
How does a model risk rating influence validation and monitoring requirements?
Ratings are frequently used to calibrate the intensity and frequency of oversight activities. Higher-rated models often warrant more rigorous validation, more frequent revalidation or monitoring, and closer review, while lower-rated models may be subject to lighter-touch processes. The precise thresholds and expectations depend on an organization's policy and any applicable guidance, so the mapping between rating tiers and required activities should be defined explicitly rather than assumed.
Who is typically responsible for assigning and reviewing model risk ratings?
Responsibilities are commonly distributed across lines of defense. Model owners or developers (often described as the first line) may propose an initial rating, while an independent function such as model risk management (frequently the second line) reviews, challenges, or approves it. Internal audit or a third line may assess whether the rating process operates as intended. The exact allocation depends on an organization's governance structure and should not be assumed to be uniform across institutions.
How often should model risk ratings be reviewed or updated?
Ratings are generally not static. Many organizations review them periodically and also reassess when material changes occur, such as changes in the model's use, data, complexity, or the significance of decisions it supports. The appropriate cadence depends on the model's rating tier and organizational policy, and there is no single universally required interval. Documenting the review triggers and frequency helps demonstrate that ratings remain current.

Common misconceptions

A model risk rating measures how well a model performs.
A model risk rating typically reflects the potential risk arising from a model's use, complexity, and materiality, which is distinct from model performance. A high-performing model can still carry a high risk rating due to its materiality or use context, and performance degradation is a separate concern from the rating itself.
Model risk ratings are defined uniformly by regulation and are comparable across organizations.
In many frameworks, rating scales, criteria, and thresholds are set internally by each institution rather than prescribed by a single authoritative standard. As a result, ratings are not necessarily interchangeable across firms, and comparisons should account for differing methodologies.
A risk rating captures the risk remaining after controls are applied.
Practitioners commonly distinguish inherent risk from residual risk. A rating may reflect inherent risk before mitigation, residual risk after controls, or both as separate values, depending on the institution's methodology. Which is being represented should be stated explicitly to avoid conflation.

Best practices

Document the specific components and criteria used to derive each rating, and clarify whether the rating reflects inherent risk, residual risk, or both, so the meaning is not ambiguous to reviewers.
Keep the rating conceptually separate from performance metrics; assess materiality, complexity, and use context independently from how accurately the model currently performs.
Define and consistently apply the rating scale and thresholds across the model inventory so that ratings support comparable prioritization within the organization.
Reassess ratings when the model's use, materiality, data, or surrounding controls change, rather than treating the rating as a one-time, static assignment.
Use the rating to calibrate the intensity of validation, monitoring, and oversight, while recognizing that a rating reduces or manages risk exposure rather than eliminating it.
State the limitations of the methodology explicitly, noting that criteria are typically institution-specific and may differ across sectors such as banking model risk versus general enterprise AI governance.