AI impact assessment severity rating is where an assessment decides how seriously to take each impact on people, and it is where many assessments become inconsistent. One assessor calls a wrongly denied application “moderate” because the applicant can reapply, another calls it “severe” because the applicant lost a home. Without a shared scale, ratings say more about the assessor than about the system.
This guide explains how to design an AI impact assessment severity rating: which factors to include, how to define levels, how to treat scale, reversibility and vulnerability, how to calibrate assessors and how to record reasons. It is general guidance that follows the spirit of ISO/IEC 42005 and should be adapted to your context.
Why severity rating matters in an AI impact assessment
An AI impact assessment looks at how a system could affect individuals, groups and society. Severity tells you which impacts need attention first and how strong the safeguards must be. Treat every impact as equally serious and resources are wasted on trivia. Underrate the serious ones and people get hurt.
ISO/IEC 42005 gives guidance on assessing impacts, including how to consider the severity and likelihood of harm, and how to involve those affected; see ISO/IEC 42005 on AI system impact assessment. The standard leaves organizations to define their own scales, which is why a clear method matters.
Start with the type of harm
Classify harms so that you consider all of them: physical or psychological harm, financial loss, discrimination and unfair treatment, loss of privacy or autonomy, denial of services or opportunities, damage to reputation and harm to society such as misinformation or erosion of trust. Our AI harm taxonomy provides a list you can adapt.
For each impact, describe who is affected and how. A precise description, such as “applicants from a particular region are more likely to be wrongly rejected”, is much easier to rate than a generic label like “bias”.
Define the scale for AI impact assessment severity rating
Use four or five levels with clear descriptions and examples, as in the table above. Define each level in terms of consequences for people, not for the organization. Include both short-term and long-term effects, and include examples relevant to your domain, such as employment, credit, healthcare or education.
Keep wording plain and concrete. “Serious” means little until you say, for instance, “loss of employment or income, denial of essential services or significant discrimination”. Test the scale on real cases and adjust wording that causes confusion.
| Level | Description | Example impact on a person |
|---|---|---|
| 1 Negligible | Minor inconvenience, easily corrected | Irrelevant product recommendation |
| 2 Limited | Temporary or minor harm, easily reversed | Short delay in a service, quickly resolved |
| 3 Significant | Notable harm, reversible with effort | Wrongful refusal of a loan corrected on appeal |
| 4 Serious | Substantial harm, hard to reverse | Loss of job opportunity or income, discrimination |
| 5 Critical | Severe or irreversible harm | Physical injury, loss of liberty, loss of home |
- Describe consequences for people, not the organization
- Include short and long-term effects
- Provide domain-specific examples
- Test the scale on real examples before use
Adjust for scale, reversibility and vulnerability
Three factors change how serious an impact is. Scale: how many people could be affected, and how often. Reversibility: can the harm be undone, and how easily? Vulnerability: are the affected people children, patients, people in financial difficulty or others with less ability to protect themselves?
Build these into the scale, or use them as modifiers. For example, raise the rating by one level where the harm is irreversible or the group is vulnerable. A modest harm to millions can be as serious as a severe harm to a few, so consider both. Compare with fundamental rights impact assessment, which looks at the same issues from a rights perspective.
Combine your AI impact assessment severity rating with likelihood
Severity is only half of the picture. Combine it with likelihood using a matrix, and consider the effect of existing safeguards. For high-severity impacts, even a low likelihood may justify strong controls, especially if harm is irreversible. Treat certain or already occurring harms as such: if testing shows a group is disadvantaged, likelihood is not hypothetical.
Record inherent and residual ratings, as with any risk method. For the wider relationship between the two assessments, see AI impact assessment vs risk assessment, and for scoring organizational risks, see AI risk scoring.
Free AI impact assessment (ISO 42005)
Who could this AI system affect, and how?
Screen the system against sensitive and prohibited uses, describe it, check the safeguards for fairness, transparency and oversight, and rate its impacts on people and society from 26 scenarios with ISO 42001 Annex A measures. Free, with findings.
Start the free AI impact assessment → or View premium report sample
Involve the people affected in the severity rating
Assessors often underrate impacts that they have not experienced. Consult representatives of affected groups, front-line staff, advocacy organizations or domain experts, and use their input to test the severity ratings. See AI impact assessment stakeholders for how to identify and engage them.
Record the views received and how they influenced the ratings. Where you disagree with the input, explain why. This creates a trail showing that severity was not decided in isolation.
Calibrate assessors and record reasons for each severity rating
Run a calibration session before the assessment. Give assessors several scenarios, ask them to rate independently and discuss the differences. Update the scale guidance with agreed interpretations. Have a second reviewer check high and critical ratings.
Record the reasoning for every rating: what harm, who is affected, how many, how reversible and what evidence supports the view. Short reasons are enough. When someone questions the rating later, the record shows how the decision was made.
Turn ratings into action
Ratings should drive decisions. Define what each level triggers: for example, levels one and two are accepted with monitoring, level three requires documented mitigation, level four needs senior approval and stronger safeguards, and level five means the system should not go live in its current form. Use the outputs in your AI risk treatment plan.
Where the residual rating remains high, consider redesign, narrower scope, more human oversight, additional testing or withdrawal. Record the decision and who took it.
Common mistakes in AI impact assessment severity rating
Frequent errors include rating impact on the organization rather than on people, using vague scale wording, ignoring reversibility and vulnerability, averaging across the population and hiding harm to small groups, skipping consultation, failing to record reasons and never recalibrating. Another is treating a low overall rating as proof of safety when a specific group is seriously harmed.
Avoid these by using concrete descriptions, disaggregating results by group, involving affected voices and reviewing calibration regularly. See the AI impact assessment example for a worked record.
Documenting severity decisions for audit
Keep the scale, its version and its approval with the assessment. For each impact, store the rating, the reasons, the evidence, the people consulted and the reviewer. If a regulator, auditor or affected person asks how a decision was reached, the file shows a consistent method rather than an ad hoc judgement. When a rating changes, keep the earlier version and note why.
Where suppliers provide the AI system, request their own impact information and severity views, and compare them with yours. Differences are worth discussing, because the supplier may see technical limitations that your team cannot, while you understand the deployment context that the supplier does not.
A short worked example
A council uses a model to prioritise home inspections. The assessors identify an impact: tenants in certain neighbourhoods are inspected less often, leaving unsafe housing unaddressed. Harm type: physical safety and housing. Scale: thousands of households. Reversibility: low, since injuries cannot be undone. Vulnerability: many tenants are on low incomes. Base severity: serious, raised to critical by irreversibility and vulnerability.
The council adds human review of low-priority areas, monitors inspection rates by neighbourhood and reports results to a committee. Residual severity is reduced to significant, and approval is given by a senior officer. The reasons are recorded, and the rating is reviewed every six months.
Keeping AI impact assessment severity rating current
Review the scale at least annually, and after incidents, regulatory changes or feedback from affected groups. Compare ratings with what actually happened: were impacts under- or over-estimated? Adjust wording and thresholds and note the change, so historic assessments can be interpreted.
Share the scale across teams, so that AI impact assessments, DPIAs and risk assessments describe harm in compatible ways. Consistency saves time and reduces the confusion that arises when each document uses different labels.
Structuring the assessment
If you want a report and workbook with screening, impacts on people, ratings, safeguards and review in one place, the AI Impact Assessment Report and Workbook provides a structured layout built around ISO/IEC 42005. Whatever tool you use, a clear AI impact assessment severity rating turns a discussion of possible harms into decisions about which ones must be prevented.
AI impact assessment severity rating FAQ
What should the severity scale measure?
Consequences for individuals and groups, such as physical, financial, psychological and rights-related harm, not just consequences for the organization.
How many levels should the scale have?
Four or five, with concrete descriptions and examples for each so different assessors reach similar ratings.
How do we treat irreversible harm?
Rate it higher or apply a modifier, since harm that cannot be undone is more serious than harm that can be corrected.
Should we rate the average or the worst-affected group?
Consider both. Average results can hide serious harm to a small or vulnerable group, so disaggregate the analysis.
How often should ratings be reviewed?
At least annually and when the system, data, users or context change, or after incidents and complaints.