Governance DocsGovernance Docs
Browse Toolkits

CART

No products in the cart.

ISO Compliance Insights & Best Practices

NIST AI RMF Measure function four categories and TEVV

NIST AI RMF Measure Function: TEVV Guide for 2026

The NIST AI RMF Measure function is where an AI risk programme stops describing risks and starts testing them, and it is the function most organisations under-build. Govern and Map produce policies and context, but without Measure, nobody can say whether a system is accurate, fair, secure or robust in the conditions it actually meets.

This guide explains what the Measure function requires, walks through its four categories, shows how test, evaluation, verification and validation fit in, and gives a practical way to put it into operation. For the framework as a whole, see our NIST AI RMF overview.

Free gap assessment

How much of ISO 42001 could you evidence today?

Score every clause and Annex A control of the AI management standard, free, and see where the programme really sits.

Run the free ISO 42001 gap assessment →  or  View premium report sample

What the NIST AI RMF Measure function does

NIST’s AI Risk Management Framework 1.0 describes Measure as the use of quantitative, qualitative or mixed-method tools, techniques and methodologies to analyse, assess, benchmark and monitor AI risk and related impacts. It takes the risks identified in the Map function and gives them numbers or structured judgments, so that the Manage function can prioritise and respond. It also draws on test, evaluation, verification and validation, usually shortened to TEVV, which the framework expects to be in place, followed and documented.

The NIST AI RMF Measure function does not stand alone. What you measure comes from Map, who reviews the results comes from Govern, and what you do about the results is Manage. Our NIST AI RMF Playbook guide lists all the subcategories across the four functions.

The four Measure categories

CategoryFocusWhat to produce
Measure 1Appropriate methods and metrics are identified and appliedMetric catalogue, list of risks that cannot be measured, independent review
Measure 2AI systems are evaluated for trustworthy characteristicsTest plans, results, evaluation reports
Measure 3Mechanisms for tracking identified AI risks over time are in placeMonitoring process, incident and feedback channels
Measure 4Feedback about the efficacy of measurement is gathered and assessedReviews of whether metrics are useful, stakeholder input

Measure 1: choose methods and metrics

Organisations select measurement approaches for the significant risks identified while mapping. The framework also expects you to document which risks or trustworthiness characteristics will not be measured, and why. That is an honest position: some harms are hard to quantify, and saying so is better than inventing a number. Regular assessments should check that the metrics remain appropriate and the controls effective. The category expects internal experts who did not serve as front-line developers of the system, or independent assessors, to be involved. That separation matters, since the builders of a system are poorly placed to judge it alone.

Measure 2: evaluate trustworthy characteristics

This is the largest category. It asks for testing to be documented, including the metrics and tools used, for human subject evaluations to meet applicable requirements, and for performance to be measured in conditions that resemble deployment. It covers the characteristics NIST associates with trustworthy AI: safety, security and resilience, transparency and explainability, privacy and fairness, and it also asks about the environmental impact and sustainability of training and managing models. Your test plan should map each characteristic to a method, a threshold and an owner.

Measure 3: track risks over time

The NIST AI RMF Measure function also covers life after launch. AI systems change after release because data drifts, users behave differently and models are updated. This category asks for mechanisms to track known, unanticipated and emerging risks during deployment, for approaches that accommodate risks that current techniques cannot measure, and for feedback processes through which end users and affected communities can report problems. Monitoring dashboards help, but so does a simple channel for users to say a result looks wrong.

Measure 4: check that measuring works

The last category closes the loop. Measurement approaches should be connected to deployment context, through consultation with domain experts. Results should be validated with relevant stakeholders, and improvements or declines in performance should be documented, with input from the affected community. In plain terms: ask whether your metrics tell you anything useful, and change them if they do not.

Turning the NIST AI RMF Measure function into a test plan

  1. Start from the risk register. Take each significant risk from your Map outputs.
  2. Link each risk to a characteristic. For example, discrimination risk maps to fairness, and data leakage maps to privacy and security.
  3. Choose a method and metric. Options include benchmark tests, red teaming, bias testing across groups, robustness testing, human review and user studies.
  4. Set thresholds. Define what result is acceptable and what triggers escalation, tied to your risk tolerance.
  5. Test in realistic conditions. Use data and scenarios that resemble deployment, not only clean benchmarks.
  6. Involve independent reviewers. Someone who did not build the system should examine the method and results.
  7. Record and report. Store the plan, data, results and decisions, and send a summary to the people who govern the system.
  8. Monitor after release. Keep testing on a schedule and after each significant change.

Measuring generative AI systems

Generative systems raise measurement problems that classic accuracy metrics do not cover. Outputs vary, harms include fabricated content, harmful instructions and leakage of training data, and testing has to include adversarial prompts as well as normal use. NIST has published a companion profile for generative AI that lists risks and suggested actions, and our generative AI profile guide explains it. In practice, combine automated evaluation sets with human review and red teaming, and repeat after each model or prompt change. Our generative AI risk assessment guide shows how to structure the risks you are measuring.

Free AI risk assessment

Which of your AI systems could harm people, or you?

List your AI systems, models and data, pick from 38 AI risk scenarios, rate them for the people affected and for you, and plan treatment with ISO 42001 Annex A controls. You get a heat map, a process score and the findings an auditor would raise, free.

Run the free AI risk assessment →  or  View premium report sample

Governance of the NIST AI RMF Measure function

Measurement needs a home. Assign a named owner for the metric catalogue, usually in a risk, model validation or responsible AI function, and make sure results reach the people accountable for each system. Set out when a test result forces a decision: a pause of a release, a fix, additional monitoring or acceptance by a named executive. Record those decisions with the data behind them. If the same team builds, tests and approves a system, the framework’s expectation of independent review is not met, so use another team, an internal audit function or an outside assessor.

Metrics that support decisions

A useful metric answers a decision. Fairness gaps between groups tell you whether to retrain, approval override rates tell you how far to trust the model, and incident counts tell you whether controls hold. Avoid a wall of numbers nobody uses. Select a small set for each system, link each to a threshold and a response, and review the set at least annually. That keeps the NIST AI RMF Measure function practical and stops it becoming a reporting burden.

A hypothetical example

A lender uses a model to pre-screen loan applications. In Map, the team identifies discrimination, data quality, and explainability risks. In Measure 1 they choose approval-rate parity checks, calibration by group and a document of the risks they will not quantify, such as long-term social effects. In Measure 2 they test on a holdout set that resembles live applicants and on stress cases, and a validation team that did not build the model reviews the method. In Measure 3 they monitor approval rates and complaint volumes weekly. In Measure 4 they meet with compliance and a consumer advocate to review whether the metrics capture what matters, and they add one. The example is illustrative only.

Common mistakes with the NIST AI RMF Measure function

  • Testing only accuracy. Fairness, security, privacy and robustness are skipped.
  • Clean benchmarks only. The tests do not resemble deployment.
  • Builders grade their own work. No independent review.
  • No thresholds. Results are recorded, but nothing triggers action.
  • One-off testing. No monitoring after release or after model changes.
  • Hidden limits. Risks that cannot be measured are not documented.

Third-party and vendor models

Many organisations do not build their models, they buy or call them. Measure still applies. Ask suppliers for their evaluation results, the data and conditions used, known limitations and update practices, and then test the system on your own data and in your own workflow, since a supplier’s benchmark may not reflect your use. Contracts should give you the right to receive evaluation information and to be told of material model changes, because a silent update can change behaviour overnight. Where the supplier will not share detail, record that as a risk, add compensating checks such as human review of outputs, and reflect it in your monitoring plan.

Documentation and templates

The records that support Measure are a metric catalogue, a test plan per system, evaluation reports, a monitoring plan, a feedback log and a summary for governance. The NIST AI RMF Toolkit includes templates for these, which you can adapt to your systems, and our NIST AI RMF templates guide shows how the four functions fit together. For the primary text, read section 5.3 of the NIST AI RMF 1.0. If you are comparing frameworks, see our ISO 42001 vs NIST AI RMF comparison.

NIST AI RMF Measure function FAQ

What is the Measure function in the NIST AI RMF?

It is the function that applies quantitative, qualitative or mixed methods to analyse, assess, benchmark and monitor AI risks and related impacts.

How many categories does Measure have?

Four: methods and metrics, evaluation of trustworthy characteristics, tracking risks over time, and feedback on the efficacy of measurement.

Who should carry out the measurement?

The framework expects internal experts who were not front-line developers, or independent assessors, to be involved in regular assessments.

What if a risk cannot be measured?

Document it, explain why, and describe the other ways you will manage or monitor it.

When a standard changes, know first

One email a month: edition changes, new deadlines, and what they mean for documentation you already have. No sales sequence.

We don’t spam! Read our privacy policy for more info.