AI Governance Institute
← Monitoring & Drift
MON · Monitoring & DriftMON-002Medium effort

Model Drift Detection

Added May 2026

Monitor live AI systems for data drift, concept drift, and shifts in outputs that signal degraded or changed model behavior.

Objective

Detect when AI model behavior has meaningfully changed from how it behaved at deployment (its baseline), triggering investigation and remediation before harm occurs.

Maturity Levels

1

Initial

No drift monitoring exists; model degradation is discovered reactively.

2

Developing

Some output metrics are monitored but changes in inputs and concept drift are not assessed.

3

Defined

Automated drift detection monitors patterns in inputs and outputs, and key performance metrics, with defined alert thresholds.

4

Managed

Drift alerts are prioritized and resolved within defined response deadlines (service level agreements, or SLAs); patterns are analyzed to identify root causes.

5

Optimizing

The drift detection tools are themselves checked for accuracy; thresholds are tuned based on how often they raise false alarms or miss real drift.

Evidence Requirements

What an auditor or assessor would expect to see for this control.

  • —Drift monitoring settings documenting methods, metrics monitored, and alert thresholds (the levels that trigger an alert) for each system
  • —Periodic drift reports statistically comparing current inputs and outputs against the baseline recorded at deployment
  • —Alert records for drift events including detection date, severity, and response timeline
  • —Root cause analysis records for significant drift events, with remediation actions and outcomes
  • —Retraining or recalibration (adjusting the model) records triggered by confirmed drift, including test results before redeployment

Implementation Notes

Key steps

  • Monitor three types of drift separately: data drift (the inputs change), concept drift (the link between inputs and correct outputs changes), and prediction drift (the outputs change).
  • For systems built on large language models (LLMs, the technology behind chatbots), standard statistical drift checks do not apply directly. Instead, monitor indirect signals: output length, topics covered, how often the system refuses requests, and human feedback.
  • Set drift alerts that notify the model owner, not just the monitoring dashboard; alerts nobody acts on are as bad as no alerts.
  • Document what actions are taken in response to drift alerts: investigation, retraining, rollback, or escalation.

Example Implementation

Product recommendation model experiencing seasonal purchasing pattern shifts

Drift Monitoring Configuration: Product Recommendation Engine

Monitored drift types and methods:

Drift TypeMethodMetricAlert ThresholdCadence
Data drift (input)Population Stability Index on top 20 featuresPSI> 0.2 (any feature)Daily
Prediction driftKS test on score distributionKS statisticp < 0.01Daily
Business outcome driftCTR on recommended productsClick-through rate> 15% drop from 7-day rolling avgHourly
Latency driftp95 inference timems> 200msReal-time

Alert routing: All alerts → #ml-monitoring Slack channel + model owner email

Response SLAs:

  • Business outcome alert: investigate within 2 hours; escalate if not resolved in 4
  • Data/prediction drift: investigate within 1 business day; present findings in weekly model review

Seasonal adjustment: Baseline updated each quarter to account for known seasonal patterns (documented in monitoring runbook)