AI Governance Institute
← AI Governance Playbook

Question 18 of 53

What is our process for model drift monitoring?

By Cody Maxwell · AI Governance Institute · February 2026 · Updated September 2026 · Last verified September 13, 2026

Assign owners and review schedules for deployed models. Monitor performance degradation, changes in behavior, and emerging bias.

▸Editorial status
AI Governance Institute recommendationVerified by AI Governance Institute pipelineNext review December 12, 2026
  • September 13, 2026 · Substantive update — The FSB consultation report explicitly requires post-deployment monitoring including drift detection and outcome fairness assessments, which directly intersects with this article's core topic. It also adds financial-sector-specific context (banks, insurers, model risk management principles, independent validation) and third-party AI oversight requirements that the current article does not address and that practitioners in regulated financial institutions would need. (AI Governance Institute pipeline)

How we verify and maintain this

If you only do 3 things, do this:

  1. 1.Assign explicit ownership for post-deployment monitoring before the system goes live. Without a named owner, monitoring gets deprioritized every time.
  2. 2.Define response thresholds in advance: how far can the model drift (its behavior or accuracy changing over time) before it is flagged for review, and how far before it is paused? Pre-set these. Don't determine them in real time when something goes wrong.
  3. 3.Track the pattern of the model's outputs over time. A model gradually approving or rejecting more of one group is drifting, even if its overall accuracy looks stable.

The Situation

Who this is for: Data science, machine learning engineering, and compliance teams responsible for AI systems in live use

When you need this: Before any AI system goes to production, or when a deployed system's behavior is called into question

The Decision

Do we have the monitoring infrastructure and governance process to detect and respond to model drift before it causes harm or regulatory exposure?

The Steps

  1. 1Define what to measure for each live system: the pattern of its outputs, its accuracy, fairness measures for each group of people affected, and the pattern of the data coming in
  2. 2Set alert thresholds for each metric: yellow (flag for review), red (pause for investigation)
  3. 3Assign a named monitoring owner for each system and a review cadence proportionate to risk
  4. 4Have your data science team set up automated monitoring with the tools they already use (for example MLflow, Evidently, or custom dashboards)
  5. 5Build escalation protocols: what triggers a pause, who makes the call, what is the response timeline
  6. 6Define re-training and re-validation criteria: when does drift require a model update vs. a retirement decision

The Artifacts

  • —Monitoring metrics specification template (by model type)
  • —Alert threshold setting worksheet
  • —Monitoring ownership register (system → owner → cadence)
  • —Drift response protocol (yellow/red thresholds → escalation path → response actions)
  • —Model performance dashboard specification
Open the implementation kit

The Output

A documented monitoring plan for every production AI system, with metrics defined, thresholds set, ownership assigned, and automated alerts in place.

Compliance does not end at deployment

A model that passes all pre-deployment testing may behave differently in production as the world changes around it. Data drift happens when the data coming into the model changes over time, for example when a new customer segment starts applying. Concept drift happens when the real-world link between the inputs and the right answer changes, often because the thing the model was trained to predict has itself changed. Both can cause a previously compliant model to become inaccurate, biased, or harmful without any change to the model itself.

Regulatory frameworks including the EU AI Act and NIST AI RMF explicitly address post-deployment monitoring requirements. For high-risk systems, the EU AI Act requires post-market monitoring plans and logging of system operation. The regulatory expectation is that governance is ongoing, not a one-time pre-deployment exercise.

What to monitor and how often

Monitor the pattern of outputs: are the model's results shifting over time? A credit scoring model that is approving an increasing proportion of applications, or a hiring model that is rejecting more candidates in a particular category, may be exhibiting drift. Standard statistical checks can flag when the pattern of outputs moves outside its expected range.

Monitor input data quality: is the data the model relies on behaving as expected? Missing values, out-of-range inputs, and shifts in key data fields are early signs of drift. Where you later learn the real outcome (whether a loan was repaid, whether a flagged transaction was fraud), compare it to what the model predicted. Track overall accuracy, how often the model's flags are wrong, how often it misses real cases, and fairness measures for each group.

Establish monitoring cadences proportionate to risk. High-risk systems in rapidly changing environments may require weekly or even daily monitoring. Lower-risk systems in stable environments may be adequately served by monthly or quarterly reviews.

Ownership and response protocols

Assign explicit ownership for post-deployment monitoring to a named role or team. Without ownership, monitoring activities are consistently deprioritized in favor of new deployments. The owner is responsible for running monitoring checks, reviewing results, escalating anomalies, and initiating retraining or retirement processes when drift is detected.

Define response thresholds in advance: at what level of detected drift does the system get flagged for review? At what level is it paused pending investigation? At what level is it retired? These thresholds should be calibrated to the risk level of the system and documented before deployment, not determined on the fly when something goes wrong.

Financial sector obligations under the FSB AI sound practices

The Financial Stability Board's consultation report on responsible AI adoption sets out 12 practices that apply to banks, insurers, and other regulated financial entities. On post-deployment monitoring specifically, the FSB expects institutions to conduct regular performance reviews that include drift detection and outcome fairness assessments, and to embed human review mechanisms at high-impact decision points. These expectations align closely with model risk management frameworks already familiar to financial firms, so institutions should map their drift monitoring programs to existing model validation and independent review processes rather than treating AI monitoring as a separate discipline.

Third-party AI systems require particular attention under the FSB guidance. If a model producing decisions is sourced from an external vendor, the institution remains responsible for ongoing monitoring and must support that obligation through contractual provisions and due diligence. Documentation requirements span the full AI lifecycle, meaning drift monitoring logs, threshold decisions, and escalation records should be retained as part of the model's lifecycle record, not discarded after each review cycle.

Turn this guidance into an implementation plan

Get the free Excel tracker for all 132 governance controls. Score maturity, assign owners, and set deadlines, including this playbook's 12 related controls.

  • 132 controls in Excel
  • Score maturity and assign owners
  • Track deadlines and regulation coverage

Includes AI Governance Weekly every Thursday. Unsubscribe anytime.