AI model drift risk is the danger that a system which worked well at launch gradually stops working well, without any obvious failure to alert anyone. Models are trained on data from the past. The world, the customers and the data feeding them keep changing. Performance decays, sometimes slowly and sometimes suddenly, and the first sign may be a complaint, a loss or an audit finding.
This guide explains what causes drift, how to assess AI model drift risk, what to monitor, how to set thresholds and what to do when thresholds are crossed.
What AI model drift risk is
Drift is a change over time in the relationship between a model and the environment it works in. It matters for risk because a drifting model can make worse decisions, treat groups unfairly or expose the organisation to legal and financial harm, while appearing to run normally. The NIST AI Risk Management Framework expects organisations to monitor deployed AI systems and to track performance and changes over time; see the framework on the NIST AI RMF page. Our guide to the NIST AI RMF Measure function covers how measurement fits into the framework.
Types of drift
| Type | What changes | Example |
|---|---|---|
| Data drift | The distribution of input data shifts | New customer segment with different income patterns |
| Concept drift | The relationship between inputs and the right answer shifts | Fraud tactics evolve so old patterns no longer signal fraud |
| Label or outcome drift | The definition or frequency of outcomes changes | A policy change alters what counts as a default |
| Upstream change | Data pipelines, sensors or vendor models change | A supplier updates a scoring model without notice |
| Feedback loops | The model’s own outputs alter future data | A recommender narrows what users see, changing what they click |
Generative and vendor-provided models add a further source: the provider may change the model or its safety settings, altering behaviour with no change on your side. Our guide to generative AI risk assessment covers those cases.
Free AI risk assessment
Which of your AI systems could harm people, or you?
List your AI systems, models and data, pick from 38 AI risk scenarios, rate them for the people affected and for you, and plan treatment with ISO 42001 Annex A controls. You get a heat map, a process score and the findings an auditor would raise, free.
Run the free AI risk assessment → or View premium report sample
Assessing AI model drift risk
Include drift as a risk in your assessment of every model in production, and score it like any other. Consider how fast the environment changes, how sensitive the model is, how quickly you would notice and how severe the consequences of wrong outputs would be. Under the approach in our guide to AI risk assessment scoring, detectability is central: a model with no monitoring has poor detectability, so its drift risk rates higher.
Factors that raise the rating include volatile domains such as fraud, markets and health, models trained on old or narrow data, high automation with little human review, dependence on third-party components, and outputs that affect people’s rights or access to services.
What to monitor
Input data
Compare the distribution of live inputs with the training data. Track missing values, ranges, category frequencies and statistical distance measures for key features. Sudden changes often point to a pipeline fault, not a real shift in the world, and both matter.
Model outputs
Watch the distribution of predictions or scores. A shift in the share of approvals, the average score or the proportion flagged may signal drift even before you have ground truth on outcomes.
Performance
Where you eventually learn the correct outcome, such as whether a loan defaulted or a transaction was fraud, calculate accuracy, error rates and calibration on recent data. Some outcomes arrive late, so use proxy measures in the meantime and plan for delayed labels.
Fairness and impact
Measure outcomes for relevant groups over time. A model may keep good overall accuracy while its error rate for one group grows. Link this to your impact assessment; see our AI impact assessment screening guide for how impact shapes monitoring depth.
Operational signals
Monitor latency, error rates, override rates by human reviewers, complaints and appeals. A rising override rate is a valuable early sign that people no longer trust the outputs.
Setting thresholds and responses
Agree thresholds in advance, using bands that tell people what to do. A green band means normal operation. An amber band triggers investigation within a set time. A red band triggers a defined response such as increased human review, rolling back to a previous model, pausing automated decisions or retraining. Record who is responsible for each step and how quickly it must happen. Base initial thresholds on validation results and expected variation, and adjust them as you learn how the model behaves.
Not every alert means the model is wrong. Investigate causes: a data pipeline fault, a seasonal effect, a genuine shift in the population or a fault in the labels. The response depends on the cause.
Retraining and change control
Retraining is the most common response, but it is a change like any other. Treat it under change control: test the new model against a holdout set and against the current model, check fairness and robustness, review documentation, get approval and plan the rollout. Retraining on data that includes the model’s own decisions can reinforce bias, so think about how labels were produced. Keep versions of models, data and code so that you can reproduce and roll back.
Vendor and third-party models
If you use a supplier’s model, you may not see the internals or receive notice of updates. Require change notification in the contract, ask for release notes and performance information, and run your own tests on representative data after each change. Where a supplier cannot tell you what changed, record it as a risk and consider a fallback. Include the supplier in your third-party assessments.
A hypothetical example of AI model drift risk
The following is a hypothetical example invented for illustration. An insurer uses a model to prioritise claims for fraud review. At launch, it flagged 4 per cent of claims and 30 per cent of reviewed flags were confirmed. Over the following year the share of flagged claims slowly falls to 2 per cent, and the confirmed rate falls to 18 per cent. No alarm sounds because the system is running and no threshold is set on either measure.
An internal audit notices the trend, and investigation finds that a change in how repair shops submit invoices altered several input features, while fraudsters adapted to earlier patterns. The insurer sets amber and red thresholds on both flag rate and confirmed rate, adds monthly drift reports to the risk committee, retrains the model with recent data and adds a rule to check the pipeline after any change in submission formats. The problem had been costing money for months.
Common mistakes with AI model drift risk
Frequent weaknesses include no monitoring after launch, monitoring only uptime, waiting for complaints, no thresholds or owners, no ground truth process, ignoring vendor updates, retraining without testing, never checking fairness over time, no rollback option and no record of monitoring in the risk register. Another is treating drift as a data science problem alone, when risk, business and operations must agree what to do when it happens.
Governance of monitoring
Monitoring only works if someone reads it. Assign named owners for each dashboard, set the review rhythm, and define what gets escalated to the risk committee. Report drift measures alongside other model risk information so leaders see a single picture. For high-impact systems, arrange independent review of the monitoring itself, to check that metrics still measure what matters and that thresholds have not been quietly relaxed. Keep the monitoring plan as documented information, and update it when the model, data or use changes.
Plan for the cost of monitoring at the start. Labels can be expensive to collect, and dashboards need maintenance. A model that cannot be monitored affordably is a model that should not run at high impact.
Recording drift in the risk register
For each model, record drift as a risk with its rating, the monitoring measures, thresholds, owners and response plans, and link it to the treatment plan. See our guides to the AI risk register and the AI risk treatment plan. Review the risk after any incident, model change or shift in use, and at a regular interval.
Templates for AI model drift risk
A structured report captures the model, the drift risks, the monitoring, the thresholds and the response in one place. The AI Risk Assessment Report and Workbook provides a report and workbook for documenting AI risks and their treatment. Whatever tool you use, apply the same structure to every model so that monitoring is consistent across the portfolio.
AI model drift risk FAQ
What is model drift?
A change over time that reduces how well a model fits its environment, for example because input data or real-world relationships change after training.
How quickly does drift happen?
It varies. Some domains such as fraud or markets change fast, others slowly. Monitoring is the only reliable way to find out for your model.
What should we monitor first?
Start with input distributions, output distributions and any performance measure you can compute, then add fairness and operational signals such as override rates.
Who owns drift risk?
The system owner is accountable, with data science responsible for monitoring, and risk or governance functions reviewing thresholds and responses.
Does drift apply to vendor models?
Yes. Vendors may change models without notice, so require change notification and test on your own data after updates.