Skip to main content
Category: Risk Assessment & Analysis

Severity Rating

Also known as: Severity Rating Scale, Severity Score
Simply put

A severity rating is a way of scoring how serious the impact of a problem or issue would be if it occurred. It typically uses a defined scale to help teams compare issues and decide which ones matter most. The exact scale and meaning of each level vary widely depending on the organization and the field in which it is used.

Formal definition

A severity rating is an ordinal measure used to evaluate the impact or seriousness of an issue based on the consequences of its effect. As commonly applied, ratings are assigned along a defined scale (for example, a numeric range or categorical labels such as mild, moderate, or severe) that organizations customize to their specific context. The evidence provided shows the term used across differing domains—such as failure mode analysis (FMEA), user-experience evaluation, and clinical or educational assessment—which indicates that both the scale structure and the criteria for each level are domain- and organization-specific rather than standardized. Note that this evidence does not establish a specific definition or scale for severity rating within AI governance or model risk management contexts, so any application to those areas should be scoped and defined locally.

Why it matters

Severity ratings give teams a shared, comparable way to express how serious a problem's impact would be, which is essential when resources are limited and issues must be prioritized. Without an agreed scale, judgments about what matters most tend to be inconsistent and difficult to defend or audit. By assigning an ordinal score to consequences, organizations can rank issues, allocate attention, and document the reasoning behind their choices.

The evidence shows the term appears across very different domains—failure mode analysis (FMEA), user-experience evaluation, and clinical or educational assessment such as language-disorder rating. This breadth is itself the key point: the scale structure and the criteria attached to each level are shaped by the field and the organization, not by a single universal standard. A severity level of "moderate" in a clinical language assessment does not carry the same meaning as a comparable label in an FMEA exercise, so borrowing a scale from one context into another without redefinition can produce misleading conclusions.

For readers working in AI governance or model risk management, it is important to note that the available evidence does not establish any specific severity-rating definition or scale for those areas. Treating a generic or borrowed scale as if it were authoritative for AI risk assessment would be a mistake; the concept is useful, but the criteria and thresholds must be defined locally and scoped to the intended use.

Who it's relevant to

Quality and Reliability Engineers
Practitioners using failure mode analysis (FMEA) apply severity ratings to score the seriousness of a failure's effect. The evidence includes a generic severity rating scale explicitly offered as a starting point to be customized for a specific organization, reflecting that FMEA severity criteria are tailored rather than fixed.
User-Experience Evaluators
In usability and UX evaluation, severity ratings are used to assess the impact or seriousness of an identified issue based on its effect on the user experience, helping teams decide which problems to address first.
Clinical and Educational Assessors
In clinical and educational settings, severity ratings characterize the seriousness of a condition—for example, coded diagnosis severity scales or language-disorder ratings that place a student in a mild, moderate, or severe range based on how far scores deviate from an expected mean. These uses show how domain-specific the criteria for each level can be.
AI Governance and Model Risk Professionals
The evidence does not establish a specific severity-rating definition or scale for AI governance or model risk management. Professionals in these fields should treat severity rating as a general concept that must be scoped and defined locally, rather than importing a scale from another domain and assuming it applies. Any thresholds or level criteria used for AI risk should be documented for the intended context.

Inside Severity Rating

Impact Dimension
A characterization of the magnitude of harm or consequence that could result from a model failure, error, or misuse. In many risk-rating schemes, severity captures the 'how bad' rather than the 'how likely,' and is assessed independently from probability or frequency.
Ordinal Scale or Tiers
Severity ratings are typically expressed on an ordinal scale (for example, low/medium/high/critical or a numeric band) rather than a precise continuous measure. The tiers are relative rankings, not exact quantities, and their meaning depends on the definitions an organization assigns to each level.
Scoping Criteria
The defined factors used to place an item into a severity tier, which vary by context. These may include affected population, financial exposure, regulatory or legal consequence, safety implications, or reputational impact. Criteria should be documented so ratings are reproducible and comparable across assessors.
Relationship to Risk (not equivalent to it)
Severity is commonly one input into an overall risk determination, often combined with likelihood to derive a risk level. Severity alone does not constitute a full risk rating; it describes consequence magnitude, which is a distinct component from the probability of occurrence.
Application Context
The domain in which the rating is applied, such as incident triage, model risk tiering, bias or fairness findings, or vulnerability management. The precise definition of severity is context- and framework-dependent, and a rating meaningful in one setting may not translate directly to another.

Common questions

Answers to the questions practitioners most commonly ask about Severity Rating.

Does a severity rating measure how likely a model problem is to occur?
No. Severity typically characterizes the magnitude or consequence of an adverse outcome if it occurs, not its probability. Likelihood is usually assessed separately, and many frameworks combine severity and likelihood to derive an overall risk rating. Treating severity as a proxy for likelihood is a common error that can distort prioritization, because a high-severity issue may be unlikely and a low-severity issue may be frequent.
Is a high severity rating the same as a high residual risk?
Not necessarily. Severity generally reflects potential impact before or independent of controls, whereas residual risk reflects the risk remaining after controls are applied. A finding can carry high inherent severity yet lower residual risk once mitigations, monitoring, or compensating controls are in place. Conflating the two can lead to either over- or under-stating the risk that actually remains.
How should we define severity levels so they are applied consistently across teams?
Many organizations define discrete severity tiers with written criteria and illustrative examples for each level, covering dimensions such as financial exposure, regulatory or legal consequence, customer harm, and operational disruption. Consistency is typically improved by anchoring each tier to observable thresholds rather than subjective judgment, and by calibrating ratings periodically across reviewers. Definitions and thresholds often vary by organization and sector, so there is no single authoritative scale.
Who should assign the severity rating for a model finding?
Practice varies, but severity is often proposed by the party identifying the issue and then reviewed or confirmed by an independent function. In institutions organized around lines of defense, model owners or developers (first line) may initiate a rating, with validation or risk management (second line) challenging and confirming it. Clear ownership and documented rationale help support consistency and auditability; the specific allocation depends on the organization's governance structure.
How do severity ratings feed into remediation timelines and escalation?
In many frameworks, higher severity ratings correspond to shorter remediation deadlines and higher levels of management or committee escalation, while lower severity items may follow standard tracking cycles. Linking severity to defined response expectations helps prioritize resources. The specific mapping between severity tiers and timelines is set by internal policy and may differ across organizations and regulatory contexts.
How often should severity ratings be revisited?
Severity ratings are commonly reassessed when relevant conditions change, such as new information about impact, changes in model use or exposure, added or removed controls, or as part of periodic review cycles. Documenting the basis for any change supports traceability. The cadence and triggers for reassessment typically depend on organizational policy rather than a universal standard.

Common misconceptions

A severity rating is the same as a risk rating.
Severity typically captures only the magnitude of potential consequence. In many frameworks a risk rating combines severity with likelihood (and sometimes other factors such as detectability). Treating severity as a complete risk measure omits the probability dimension and can misrank items.
Severity levels mean the same thing across every framework, tool, or organization.
Severity scales are usually ordinal and defined locally. A 'high' in one incident-management scheme, one model risk tiering approach, or one fairness-assessment method may reflect different criteria and thresholds. Ratings are not automatically interchangeable without mapping the underlying definitions.
A high severity rating means the harm will occur or that controls have failed.
Severity describes potential consequence magnitude, not the current state or an outcome. A high-severity item may have a low likelihood or be well controlled. Severity assessment reduces uncertainty about consequence but does not on its own indicate that a failure has happened or that risk has been eliminated.

Best practices

Document explicit, reproducible criteria for each severity tier so that different assessors assign consistent ratings, and record which factors (for example affected population, financial exposure, regulatory consequence) drive each level.
Keep severity distinct from likelihood in your methodology, and be transparent about how the two are combined if you derive an overall risk level from them.
Define severity within a stated scope and application context, and avoid reusing a rating produced under one scheme in a different setting without mapping the definitions.
Treat ordinal severity tiers as relative rankings rather than precise measures, and avoid implying quantitative precision the scale does not support.
Periodically review and calibrate tier definitions and sample ratings to guard against drift, inconsistent interpretation, or accumulation of items in a single tier.
Communicate severity ratings as measures that inform prioritization and risk management, not as guarantees that harm will or will not occur or that residual risk has been removed.