Measuring fairness — and governing it

Lesson 5 of 5 in Generative AI Risks and Bias: NIST-AI-600-1 and SP 1270.

Diagnosis needs measurement, and fairness measurement has a property every governance professional must internalize: “fair” is not one number — it is a family of mutually incompatible numbers. Three metric families dominate practice, and choosing among them is a values decision wearing a math costume.

The three fairness-metric families
Metric familyWhat it equalizesFits best when…Blind spot

Demographic parity

Selection rates across groups (the logic behind the four-fifths rule and NYC LL144’s impact ratios)

Gatekeeping contexts where historical access was the problem — hiring pipelines, ad delivery, school admissions

Ignores qualification differences entirely; can be “satisfied” by selecting randomly within a group

Equalized odds

Error rates — false positives and false negatives — across groups

High-stakes classification where mistakes are the harm: fraud flags, recidivism scores, medical alerts

Can require different score thresholds per group, which some read as disparate treatment — a genuine legal tension

Calibration

Score meaning: a “70% risk” means 70% for every group

Scores handed to human decision-makers who must be able to trust them uniformly

A calibrated model can still produce very different false-positive burdens across groups — exactly the COMPAS fight

One more measurement trap before the governance turn: removing protected attributes does not remove bias. ZIP code, shopping patterns, phone model, even typing cadence reconstruct race, sex, and health status statistically — the proxy problem, or redlining by algorithm. Illinois HB 3773 (effective January 2026) wrote the lesson directly into statute, expressly prohibiting employers’ use of ZIP code as a proxy for protected classes. Blindness is not fairness; only disaggregated outcome testing — measuring results by group, which requires knowing the groups — can show whether a system discriminates.

SP 1270’s governance prescriptions read, in hindsight, like a draft of the state laws that followed:

  • Impact assessments before deployment → Colorado’s deployer impact assessments (original SB 24-205) and the CCPA risk-assessment regulations.
  • Independent testing and monitoring → NYC Local Law 144’s annual independent bias audits with published impact ratios.
  • Diverse teams and participatory design → GOVERN 3 and the stakeholder-consultation language recurring in state bills.
  • Challenge and recourse processes → the human-review and appeal rights in Colorado SB 26-189 and the CCPA ADMT rules.

That lineage is the reason to study a 2022 technical publication in 2026: SP 1270 is where enforcers and legislators learned their bias vocabulary. EEOC disparate-impact thinking about hiring algorithms, FTC unfairness theories about reckless deployment, and the audit statutes all draw on its framing. When a regulator asks how you manage bias, the three-family answer is the fluent answer.

Key terms: demographic parity, equalized odds, calibration, disparate impact, impact assessment

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.