Governance DocsGovernance Docs
Browse Toolkits

CART

No products in the cart.

ISO Compliance Insights & Best Practices

AI bias testing workflow from defining groups to metrics, thresholds, remediation and monitoring

AI Bias Testing Guide 2026: Methods and Records

AI bias testing is the structured examination of whether an AI system produces systematically different results for different groups of people in ways that are unfair, unlawful or contrary to your objectives. It is one of the most concrete controls in an AI risk program, and one of the easiest to do badly, by testing on the wrong data, choosing a metric that hides the problem or running the test once and never again.

This guide explains how to design AI bias testing, which metrics to consider, how to set thresholds, how to investigate causes and fix them, and what records to keep so that your work can be reviewed by auditors, regulators and the people affected.

Why AI bias testing is part of risk assessment

Bias is a source of harm to individuals, and also of legal, financial and reputational risk to the organization. The NIST AI Risk Management Framework lists fairness, with harmful bias managed, among the characteristics of trustworthy AI, and describes several types of bias, including statistical, systemic and human. You can read the framework on the NIST AI Risk Management Framework page. The EU AI Act requires providers of high-risk systems to examine training, validation and testing data for possible biases likely to affect health, safety or fundamental rights or lead to prohibited discrimination. Testing produces the evidence that these expectations are met.

Testing findings belong in your AI risk register, rated with the same scales as other risks, and they should inform the AI harm taxonomy categories you use for fairness harms.

Free AI risk assessment

Which of your AI systems could harm people, or you?

List your AI systems, models and data, pick from 38 AI risk scenarios, rate them for the people affected and for you, and plan treatment with ISO 42001 Annex A controls. You get a heat map, a process score and the findings an auditor would raise, free.

Run the free AI risk assessment →  or  View premium report sample

Step one of AI bias testing: define what fair means for this system

There is no single definition of fairness that fits every use. Before choosing a metric, decide with legal, business and affected stakeholders what outcome matters. For a hiring tool, it may be equal selection rates across groups. For a medical triage aid, it may be equal error rates, so that no group is more likely to be missed. For a credit model, law and regulation may set specific expectations. Record the decision and reasons, because different fairness metrics can conflict, and it is generally impossible to satisfy all of them at once.

Choose the groups to compare

List the characteristics that are protected under the laws that apply, such as sex, age, race, ethnicity, disability and religion, and any others relevant to your context, such as language, region or income band. Consider intersections, for instance older women or younger disabled people, since a system can appear fair across each characteristic separately and still disadvantage a combination. Note where data on these characteristics is not held, and how you will handle that lawfully, for example by using self-reported data collected for monitoring, or by careful proxies with documented limits.

Metrics used in AI bias testing

MetricWhat it comparesTypical use
Selection rate ratioShare of each group receiving a favorable outcomeHiring, promotion, access to services
Demographic parity differenceGap in positive outcome rates between groupsScreening and ranking tools
Equal opportunity differenceGap in true positive rates for those who qualifyLending, admissions, triage
Equalized oddsGaps in both true and false positive ratesRisk scoring, fraud, safety uses
Calibration by groupWhether scores mean the same thing for each groupPredictive scores used for decisions
Error rate parityDifferences in false negative or false positive ratesBiometric or classification systems

In United States employment practice, the four-fifths rule of thumb from the Uniform Guidelines on Employee Selection Procedures treats a selection rate for a group of less than 80 percent of the rate for the group with the highest rate as evidence of possible adverse impact. It is a screening guide, not a legal safe harbor, and other jurisdictions apply different tests. New York City’s Local Law 144 requires an independent bias audit of automated employment decision tools before use, with published results. Check the rules that apply where you operate.

Building test data and running the tests

Test on data that resembles your real population, with enough examples in each group to give meaningful results. Small groups produce noisy estimates, so report confidence intervals and avoid drawing conclusions from a handful of cases. Use held-out data that was not used for training, and where possible test with real outcomes as well as synthetic examples. For language and image systems, build test sets that vary names, accents, dialects, skin tones, ages and contexts, and check for stereotypes in generated content. Our guide to generative AI risk assessment explains further checks for such systems.

  1. Prepare data. Document sources, group labels, sample sizes and known gaps.
  2. Run baseline tests. Compute chosen metrics for each group and intersections with enough data.
  3. Compare with thresholds. Flag results outside the agreed limits.
  4. Investigate causes. Check data, features, labels, thresholds and deployment context.
  5. Remediate and retest. Apply fixes and confirm the effect on both fairness and accuracy.
  6. Document and approve. Record results, decisions and residual differences.

Setting thresholds and deciding what to do

Define in advance what size of gap triggers action, and who decides. Thresholds depend on the stakes: the tolerance for a gap in a music recommendation is very different from that for a loan decision. Write the rules into your risk criteria so that testers are not bargaining over each result. When a test breaches the threshold, the options are to fix the system, to limit its use, to add human review or to withdraw it. Where a residual difference remains and is judged justified, record the justification and the approver.

Finding and fixing the causes

Bias often comes from the data: under-representation of some groups, labels that reflect past discrimination, features that act as proxies for protected characteristics or measurement that is less accurate for some people. It can also come from design choices, such as an objective that rewards the wrong outcome, thresholds set on the majority group, or deployment in a context the model was not built for. Techniques to reduce bias include improving and rebalancing data, removing or constraining proxy features, adjusting thresholds by group where lawful, retraining with fairness constraints and adding human review of borderline cases. Every fix should be retested, because fixing one metric can change another. Our article on AI model drift explains why testing must continue after launch.

Continuous monitoring after deployment

A model that passes at launch can develop gaps as data and users change. Monitor outcomes by group in production, at a frequency proportionate to the risk, with alerts when metrics cross thresholds. Track complaints and overrides by group. Retest after every model update, new data source or change of use. Feed all monitoring results into the AI risk treatment plan so that actions have owners and dates.

A short worked example

A company tests a model that ranks support tickets. Results show that tickets from customers writing in a second language are ranked lower, and the difference in average priority is beyond the threshold set in advance. Investigation finds that the model learned to treat unusual grammar as a signal of low urgency. The team adds diverse phrasing to the training data, removes a feature that captured writing style, retests and finds the gap has fallen within the limit with no loss in accuracy. It records the test data, metrics, cause, fix, retest results and the approver, and sets a quarterly monitoring check on priority by language group.

Third-party and vendor models

If a supplier provides the model, ask what testing it performed and on which populations, and run your own tests on your own data, since bias depends on the population and context. Include the right to receive updated results in the contract. Our guide to third-party AI impact assessment covers what to request and how to handle gaps in supplier information.

Records to keep

Keep the definition of fairness chosen and its rationale, the groups and how labels were obtained, the test data description, the metrics and thresholds, the results with confidence intervals, the investigation of causes, the changes made, the retest results, the decision and approver, and the monitoring plan. Store the tests so they can be rerun, including code and data versions. These records are what an auditor or regulator will ask to see, and they let a new team member understand what was done.

Common mistakes in AI bias testing

Teams test only on the training data, use a single metric without explaining why, ignore intersections, leave out groups with small samples without saying so, run the test once at launch, treat a passing score as proof of fairness and fail to record decisions. Another mistake is testing for bias only after the model is finished. Build the tests into design so that findings can change the model, not just the paperwork.

Using a ready structure

If you would like a starting structure for your assessment, the AI Risk Assessment Report and Workbook provides a structured report with scoring and a working register in which bias findings, thresholds and treatments can be recorded. Whichever format you use, AI bias testing should be planned, repeatable and documented.

AI bias testing FAQ

What is AI bias testing?

It is the systematic checking of whether an AI system’s outputs or errors differ across groups in ways that are unfair or unlawful, using defined metrics, thresholds and records.

Which fairness metric should I use?

It depends on the use and the applicable law. Different metrics can conflict, so decide with legal and business input what fairness means for the system and record the reasons.

Do I need protected characteristic data to test?

You usually need some way to identify groups. Collect self-reported data lawfully for monitoring where possible, or use documented proxies with their limits stated.

How often should testing be repeated?

Repeat it after each model update, data change or change of use, and monitor continuously in production at a frequency that matches the risk.

What if a test shows a gap?

Investigate the cause, fix or limit the system, add human review where appropriate, retest and record the decision. Any accepted residual difference needs a documented justification and approver.

When a standard changes, know first

One email a month: edition changes, new deadlines, and what they mean for documentation you already have. No sales sequence.

We don’t spam! Read our privacy policy for more info.