Governance DocsGovernance Docs
Browse Toolkits

CART

No products in the cart.

ISO Compliance Insights & Best Practices

AI risk scoring matrix combining likelihood, severity, autonomy and scale into low, medium, high and critical ratings

AI Risk Scoring: A Practical 2026 Guide

AI risk scoring gives an organization a consistent way to compare very different AI risks, from a biased hiring model to a chatbot that leaks data, and to decide which ones need attention first. Without a shared method, ratings reflect who shouts loudest, and the risk register becomes a list of opinions.

This guide explains how to design AI risk scoring: what to rate, how to define likelihood and severity for AI, which modifiers such as autonomy and scale to add, how to set thresholds and how to turn scores into treatments. It is general guidance that you should adapt to your risk framework and to the regulations that apply.

Why AI risk scoring needs its own care

AI risks share the basic shape of other risks: something could go wrong, with some probability and some consequence. But AI adds features that make scoring harder. Failures can be systematic rather than random, harms can fall on groups, behaviour can change as data changes and models can be opaque, so the causes are hard to see.

A good scoring method acknowledges these features. It uses the same overall structure as your enterprise risk framework, so results can be combined, but it adds guidance and modifiers for AI. The NIST AI Risk Management Framework, described in the NIST AI Risk Management Framework, organises the work into govern, map, measure and manage functions, and scoring sits in the measure and manage stages.

Decide what you are scoring

Score a specific risk in a specific system: “the hiring model ranks women lower for technical roles” rather than “AI bias”. Each risk should have a cause, an event and a consequence, and a clear owner. Group related risks so the register is manageable.

Start from your AI model inventory, so you know which systems exist, and from a harm taxonomy such as the AI harm taxonomy, so you consider harms to individuals, groups, the organization and society. Then record each risk in the AI risk register.

Free AI risk assessment

Which of your AI systems could harm people, or you?

List your AI systems, models and data, pick from 38 AI risk scenarios, rate them for the people affected and for you, and plan treatment with ISO 42001 Annex A controls. You get a heat map, a process score and the findings an auditor would raise, free.

Run the free AI risk assessment →  or  View premium report sample

Rate likelihood in AI risk scoring

Define likelihood levels in plain terms, such as rare, possible, likely and almost certain, with a rough frequency or evidence guide for each. Base the rating on evidence: test results, incident history, known failure modes, quality of training data, monitoring coverage and the strength of controls.

Remember that some AI failures are not rare events but constant tendencies. If testing shows a fairness gap on every run, the likelihood is not “possible”; the harm is occurring. Use testing outputs, for example from AI bias testing, and drift monitoring, such as in AI model drift, to ground the rating.

MinorModerateMajorSevere
RareLowLowMediumMedium
PossibleLowMediumMediumHigh
LikelyMediumMediumHighHigh
Almost certainMediumHighHighCritical

Rate severity for AI

Severity asks how bad the consequence would be. Consider harm to individuals (financial, physical, psychological, discriminatory), to groups, to the organization (legal, financial, reputational) and to society. Rate the worst credible consequence at each level of your scale, with concrete examples.

Severity depends on the domain. A wrong film recommendation is minor. A wrong medical triage or a wrongful benefit denial can be severe. Consider reversibility: harms that cannot be undone deserve a higher rating. Consider affected populations: vulnerable people raise severity.

  • Individuals: discrimination, loss, injury, loss of rights
  • Groups and society: unfair outcomes, misinformation, erosion of trust
  • Organization: legal, financial, operational and reputational impact
  • Reversibility and number of people affected

Add AI-specific modifiers

Many teams add modifiers that adjust the base score for factors that increase exposure. Useful ones are level of autonomy (does a person review outputs?), scale (how many people or decisions), opacity (can we explain outputs?), data sensitivity, degree of reliance by users and the maturity of monitoring.

Keep modifiers simple, for example adding one level to the rating when a system makes fully automated decisions about individuals at scale. Explain each modifier in guidance so that assessors apply them consistently. Do not add so many that the score becomes an unexplainable calculation.

Score inherent and residual risk

Score the risk before additional controls to see the exposure, then re-score after current controls to see what remains. Controls include human oversight, testing, monitoring, guardrails, training and contractual protection. See human oversight of AI for one of the most important.

Only credit controls that are in place and working. A control planned for next quarter should not reduce the score today. Record the evidence for each control, and revisit the score when evidence changes.

Decide what each rating level means: for example low may be accepted by the system owner, medium needs a documented treatment, high needs senior approval and a dated plan, and critical means the system should not go live or must be paused. Tie thresholds to your risk appetite.

Then link scores to the AI risk treatment plan: each risk above the threshold needs an owner, actions and dates. Report the distribution of scores to leaders, and highlight risks that exceed appetite.

Calibrate and review AI risk scoring

Different assessors will rate differently, so calibrate. Run sessions where assessors score the same examples, discuss the differences and refine the guidance. Have a second person review high and critical ratings. Track how scores change over time and compare them with real incidents to see whether the method is working.

Review scores whenever the system, data, use case or context changes, and at least once a year. Include review triggers such as drift alerts, new regulation and incidents. Scores are snapshots, not permanent labels.

Common mistakes in AI risk scoring

Frequent errors include scoring generic risks rather than specific ones, ignoring evidence from testing, rating fairness problems as rare when they are constant, giving credit for planned controls, ignoring scale and autonomy, using overly complicated formulas and never reviewing the scores. Another is treating the number as truth rather than a structured judgement.

Avoid these by using clear definitions, evidence-based ratings, a small number of modifiers and regular calibration. Compare with the approach in ISO 42001 risk assessment if you are working towards certification.

Scoring generative AI and third-party models

Generative systems and supplier models add uncertainty because you may not control the training data or the model updates. Score the risks you can observe, such as harmful outputs, data leakage and over-reliance, and treat unknowns conservatively. For supplier models, ask for testing evidence, model cards and change notifications, and reflect gaps in the score. The approach in generative AI risk assessment helps you identify the risks to score.

Because supplier models can change without notice, add a review trigger for every version change. A score that was fair last quarter may not hold after an update, so keep monitoring evidence close to the rating and revisit it when the behaviour of the system shifts in ways your testing did not anticipate.

A short worked example

A company uses a model to prioritise customer support tickets. A risk is identified: the model deprioritises tickets written in non-standard English, delaying help for some customers. Testing shows a consistent gap, so likelihood is likely. Consequences are moderate for most, but severe for customers reporting urgent safety issues, so severity is major. The base rating is high.

The system is autonomous and operates at scale, which raises the score to critical before controls. Adding a language-robustness test, a human escalation rule for safety keywords and monthly monitoring reduces likelihood and severity, giving a residual rating of medium. The owner records the plan, and the score is reviewed quarterly.

Documenting the method

Write a short scoring guide: definitions of likelihood and severity with examples, the modifier rules, the thresholds and the review rules. Approve it at a governance forum and keep versions. Use it in training, so assessors and system owners apply it the same way. When auditors or regulators ask how you rate AI risks, you can show the method, the calibration and the results. Update the guide when regulation, technology or your experience changes the picture, and note the change so historical scores remain interpretable.

Finally, keep a record of the scores and the reasoning for each system. That record shows how decisions were made and lets you learn from incidents and near misses.

Structuring the assessment

If you want a report and workbook that carry system descriptions, risks, scores, controls and treatments together, the AI Risk Assessment Report and Workbook provides a structured layout for AI risk assessment. Whichever tool you use, sound AI risk scoring is specific, evidence-based and consistently applied, and it leads to decisions.

AI risk scoring FAQ

Can we use our normal risk matrix for AI?

Yes, for the basic structure, which lets you combine AI risks with other risks. Add AI-specific guidance and modifiers for autonomy, scale, opacity and data sensitivity.

How do we rate likelihood for a bias risk?

Use testing evidence. If tests show a consistent gap, the harm is already occurring and likelihood should be rated accordingly, not treated as a rare event.

Should scores change when controls are added?

Yes, but only for controls that are in place and evidenced. Record inherent and residual scores separately.

Who approves high AI risks?

A senior person with authority over the system and the risk appetite, with advice from risk, legal and privacy functions.

How often should scores be reviewed?

At least annually and whenever the system, data, use case, regulation or monitoring results change.

When a standard changes, know first

One email a month: edition changes, new deadlines, and what they mean for documentation you already have. No sales sequence.

We don’t spam! Read our privacy policy for more info.