AI Governance Institute
← AI Governance Playbook

Question 13 of 53

How do we measure and mitigate algorithmic bias?

By Cody Maxwell · AI Governance Institute · February 2026 · Updated October 2026 · Last verified October 1, 2026

Choose metrics for testing discrimination against protected groups. Define what happens when testing finds bias.

▸Editorial status
AI Governance Institute recommendationVerified by Cody MaxwellNext review December 30, 2026
  • October 1, 2026 · Correction — Correction: the NYC automated employment decision tool law is Local Law 144, not Local Law 27. (Cody Maxwell)

How we verify and maintain this

If you only do 3 things, do this:

  1. 1.Choose your fairness definition before you start testing. Common definitions include equal approval rates across groups (demographic parity), equal error rates across groups (equalized odds), and treating similar people alike (individual fairness), and they cannot all be satisfied at once. This is a values decision, not a technical one.
  2. 2.Run the four-fifths rule as a first-pass screen on every system that makes decisions about people. If any group's selection rate is below 80% of the top group, investigate further.
  3. 3.Document before-and-after metrics for every remediation step. Without these records you can't defend your testing if challenged.

The Situation

Who this is for: Data science, machine learning engineering, and compliance teams responsible for fairness testing

When you need this: Before deploying any AI system that makes decisions about individuals, and periodically after deployment

The Decision

Are we measuring the right kind of fairness for our specific use case, and are our remediation records defensible?

The Steps

  1. 1Define the fairness criteria for the specific use case (demographic parity, equalized odds, individual fairness, or a combination)
  2. 2Identify all relevant protected characteristics and intersectional combinations to test
  3. 3Run the four-fifths rule as a screening test on selection rates by group
  4. 4For yes/no systems, compare error rates for each group (wrongly approved and wrongly rejected); for systems that produce scores, compare how scores are spread across groups
  5. 5If bias is detected, diagnose the source: training data, proxy variables, or model design
  6. 6Select and apply remediation technique(s); re-test; document before-and-after metrics
  7. 7Schedule periodic re-testing on the same cadence as model reviews

The Artifacts

  • —Fairness definition selection worksheet (use case criteria → recommended fairness metric)
  • —Disparate impact analysis template (selection rates, four-fifths calculator)
  • —Error rates by group template (wrongly approved and wrongly rejected)
  • —Bias remediation decision tree (bias source → technique options)
  • —Bias testing audit trail (findings, remediation, before/after metrics, re-test date)
Open the implementation kit

The Output

Documented fairness definitions for every decision-making AI system, pre-deployment bias assessments on file, and a re-testing schedule aligned with model review cadence.

Defining bias in measurable terms

Algorithmic bias is not a single phenomenon. It can manifest as disparate treatment (the model uses a protected characteristic as an input), disparate impact (the model produces systematically different outcomes for protected groups even without using the characteristic directly), or intersectional bias (the model performs differently for individuals who belong to multiple protected groups simultaneously).

Before testing, define what fairness means for your specific use case. Several mathematically precise fairness definitions exist, and they are often mutually incompatible. Demographic parity requires equal selection rates across groups. Equalized odds requires that the system be equally accurate for each group: it should correctly approve qualified people, and wrongly approve unqualified ones, at the same rates across groups. Individual fairness requires that similar individuals be treated similarly. Choosing a definition involves value judgments about what kind of error is most harmful, and that choice should be made explicitly, not left to default.

Standardized testing metrics

The four-fifths rule (also called the 80% rule) from the Equal Employment Opportunity Commission's (EEOC) Uniform Guidelines provides a widely accepted starting point for employment contexts: if the selection rate for any group is less than 80% of the rate for the most selected group, adverse impact is indicated. This is a screening tool, not a legal standard, but it provides a defensible threshold for triggering further investigation.

For models that make yes/no decisions, measure errors separately for each protected group. Compare how often each group is wrongly approved, how often each is wrongly rejected, and overall accuracy. Significant differences in error rates across groups indicate bias that may produce discriminatory outcomes even if overall accuracy is high. For scoring models used in credit or risk assessment, compare score distributions and cutoff outcomes across groups.

Remediation and documentation

Bias remediation options depend on where the bias originates. Bias in the training data may be addressed by rebalancing it: adding examples, removing duplicates, or giving underrepresented groups more weight. Other techniques build fairness rules into the model while it is being trained. A third option adjusts the model's outputs after the fact to meet the chosen fairness definition.

Document every bias finding and every remediation step. Record the metrics before and after remediation, the technique applied, and the rationale for choosing it. Bias testing and remediation records should be retained as part of the model's audit trail and reviewed whenever the model or its deployment context changes materially.

Turn this guidance into an implementation plan

Get the free Excel tracker for all 132 governance controls. Score maturity, assign owners, and set deadlines, including this playbook's 3 related controls.

  • 132 controls in Excel
  • Score maturity and assign owners
  • Track deadlines and regulation coverage

Includes AI Governance Weekly every Thursday. Unsubscribe anytime.