Governance DocsGovernance Docs
Browse Toolkits

CART

No products in the cart.

ISO Compliance Insights & Best Practices

AI risk assessment scoring matrix with likelihood impact and detectability

AI Risk Assessment Scoring: A Practical Method 2026

AI risk assessment scoring is where good intentions about responsible AI meet the need to make decisions. A risk register full of unranked concerns cannot tell a board which system to fix first or which model may go live. Scoring gives you a shared, repeatable way to compare risks across chatbots, scoring models and vendor tools, and to link each score to an action.

This guide sets out a practical approach to AI risk assessment scoring: which factors to score, how to define the scales, how to calibrate them, and how to turn a score into a decision.

What AI risk assessment scoring should measure

The NIST AI Risk Management Framework describes risk as a measure of the probability that an event occurs and the magnitude of its consequences. That gives two basic factors, likelihood and impact. For AI systems, a third factor is worth adding: detectability, or how likely you are to notice the problem before it causes harm. A model that degrades quietly is more dangerous than one that fails loudly, so a scoring approach that ignores detection understates real exposure. You can read the framework itself on the NIST AI RMF page.

FactorQuestionExample evidence
LikelihoodHow probable is the failure or misuse?Test results, incident history, exposure to adversarial input
ImpactHow serious are the consequences for people and the organisation?Affected population, reversibility, legal exposure
DetectabilityHow likely is it to be caught before harm occurs?Monitoring, human review, alerting

Impact should cover harm to individuals and society as well as business loss. Our guide on AI impact assessment vs risk assessment explains why the two views differ and how to keep them consistent.

Free AI risk assessment

Which of your AI systems could harm people, or you?

List your AI systems, models and data, pick from 38 AI risk scenarios, rate them for the people affected and for you, and plan treatment with ISO 42001 Annex A controls. You get a heat map, a process score and the findings an auditor would raise, free.

Run the free AI risk assessment →  or  View premium report sample

Defining scales for AI risk assessment scoring

Every scale needs written definitions, otherwise two assessors will score the same risk differently. Use a five-point or four-point scale and describe each level in plain terms tied to your own context.

A likelihood scale

Describe levels using frequency or evidence rather than adjectives alone. For example: rare means no plausible path or no history and strong controls; possible means it could happen and similar systems have had incidents; likely means it has happened in this or comparable systems or testing shows a repeatable weakness. Anchor levels to something you can check.

An impact scale

Include several dimensions: harm to individuals, regulatory or legal consequences, financial loss, service disruption and reputational damage. Take the highest rating across dimensions rather than averaging, because a severe harm to individuals is not offset by low financial cost. Include reversibility, since an outcome that cannot be undone deserves a higher rating.

A detectability scale

Define levels by how failure would be noticed: automatically within minutes, by routine review within days, only through complaints, or not at all. This mirrors the detection rating used in failure mode analysis, and it forces a conversation about monitoring that many AI teams have not had.

Calculating an AI risk assessment scoring result

A simple approach multiplies likelihood by impact to give a score, then adjusts for detectability. Some organisations multiply all three; others use detectability to move a risk up or down a band. Either is acceptable if it is written down and applied consistently. Multiplying all three on five-point scales gives a range from 1 to 125, which is precise-looking but rarely as accurate as it appears, so map the results to three or four bands rather than arguing over single points.

Whatever formula you use, guard against false precision. A score of 48 versus 52 does not mean much when the inputs are judgments. Bands such as low, medium, high and critical tell decision makers what they need to know.

Linking AI risk assessment scoring to decisions

A score matters only if it changes what happens. Define in advance what each band means.

  • Low. Accept and monitor through normal review.
  • Medium. Add treatment actions with owners and dates; review at the next cycle.
  • High. Senior approval to proceed, with a funded treatment plan and tighter monitoring.
  • Critical. Do not deploy or suspend use until the risk is reduced.

The treatment side is covered in our guide to the AI risk treatment plan, and the register that holds the scores is described in the AI risk register guide. Scores should appear in the register with the date, the assessor, the rationale and the treatment status.

Calibrating AI risk assessment scoring across teams

Scales drift when different teams interpret them differently. Run a short calibration session: give assessors five example risks, have them score independently, compare results and resolve differences by refining the definitions. Repeat when new assessors join. Keep the examples as a reference set so future assessors can compare their judgment with a known result.

A second check is to review the distribution. If nearly every risk is scored medium, the scales are not discriminating. If every risk sits at the top, the scales are too severe or the assessors are being defensive. A healthy register shows a spread, with a small number of risks at the top that genuinely drive attention.

Documenting the rationale

Every score should carry a sentence explaining why. “Likelihood rated likely because bias testing in March showed a repeatable gap” is far more useful than a bare number, because a later reviewer can check whether the reasoning still holds. Record the sources of evidence, such as test reports, incident logs and supplier documents, and the names of the people consulted. This also makes it much easier to defend a score to an auditor, a regulator or a challenging executive.

Scoring before and after controls

Record an inherent score, the score before controls, and a residual score, after controls. The gap shows what your controls achieve and prevents you from claiming credit for measures that are not yet working. For AI, credit for controls such as human review should depend on whether the reviewers actually have time, information and authority to override the system. A person who approves every output without reading it does not lower risk. Testing and monitoring evidence, such as the measurement practices described in our guide to the NIST AI RMF Measure function, supports the residual score.

A hypothetical example of AI risk assessment scoring

The following is a hypothetical example invented for illustration. A lender assesses a model that recommends credit limits. One risk is that the model performs worse for a particular customer group. Likelihood is rated likely, because testing showed a gap; impact is high, because the effect is on access to credit and is hard for customers to see or challenge; detectability is low, because no monitoring by group exists. The combined band is critical.

The treatment plan adds fairness testing before each release, monitoring of outcomes by group and a route for customers to ask for review. After controls, likelihood falls to possible, detectability improves to medium, and the residual band is medium. The model proceeds with the monitoring as a condition and a review after three months. The score changed the decision because it linked directly to the action taken.

Common mistakes in AI risk assessment scoring

Recurring weaknesses include scales with no written definitions, scoring only technical risks, ignoring detectability, averaging impact dimensions, treating a number as more accurate than it is, scoring once at launch and never again, and letting the system owner score alone. A further problem is scoring vendor systems by asking the vendor, when the deployer’s own use case determines the impact. Ask what the tool does in your context and who is affected.

Keeping AI risk assessment scoring current

Re-score when the model, data, use case, user group or law changes, and after any incident. Track how many risks were rescored, how many moved bands and why. Generative systems in particular change with each provider update, so include the vendor’s release notes as a trigger. Our guide to generative AI risk assessment covers the extra risks those systems bring.

A structured report for AI risk assessment scoring

A consistent format lets you compare systems and show auditors how scores were reached. The AI Risk Assessment Report and Workbook provides a report and workbook for documenting AI risks, scores, controls and residual ratings in one place. Whatever format you use, keep the same scales and fields across all assessments.

AI risk assessment scoring FAQ

Do we need a numeric score for every AI risk?

No. Bands with clear definitions are enough, provided they are applied consistently and tied to actions. A number helps only when the scales behind it are defined.

Should detectability be part of the score?

It is worth including. AI failures such as drift or bias can go unnoticed for long periods, so a risk that is hard to detect deserves more attention than the same risk with strong monitoring.

Who should score AI risks?

The system owner should propose scores, with review by risk, legal, privacy or security staff. Independent challenge improves consistency and reduces optimism.

How often should scores be reviewed?

At a regular interval, such as annually or more often for high risks, and whenever the model, data, use case or law changes or an incident occurs.

Can we reuse a risk score from a vendor?

Treat it as an input only. Your context, data and affected people determine the impact, so you should score the risk for your own use.

When a standard changes, know first

One email a month: edition changes, new deadlines, and what they mean for documentation you already have. No sales sequence.

We don’t spam! Read our privacy policy for more info.