Running the LL144-grade bias audit

Lesson 3 of 5 in Capstone: Building and Defending a US AI Compliance Program.

Week five: SiftIQ needs its bias audit before any Manhattan placement. Treat the LL144 audit as the template for outcome testing everywhere — done once, properly, it feeds the Illinois defense file, the Colorado fault-allocation record, the CCPA risk assessment, and the EEOC posture all at once (MEASURE 2, rendered five ways).

Scoping first. Is SiftIQ an AEDT at all? LL144 covers tools that substantially assist or replace discretionary decisions. The temptation — indulged by most of the market, which is why observed compliance stayed embarrassingly low — is to declare that humans “really decide” and skip the audit. Resist it: recruiters work SiftIQ’s ranked queue top-down, and rank determines who exists to the recruiter. That is substantial assistance. Document the scoping conclusion either way; an unexplained out-of-scope call is the first thing a DCWP inquiry will test.

Data next. Use historical data from your own use where you have it; the rules permit test data when history is insufficient — but the audit must say which was used and why. An audit on the vendor’s polished test set, when twelve months of Windrose production history exists, is an audit built to be impeached.

SiftIQ audit excerpt — selection rates and impact ratios (advance-to-interview decision)
CategoryApplicantsSelectedSelection rateImpact ratio (vs highest)

Men

4,200

1,470

35.0%

1.00

Women

3,800

1,178

31.0%

0.89

White

3,900

1,404

36.0%

1.00

Black

1,850

536

29.0%

0.81

Hispanic

1,600

480

30.0%

0.83

Asian

1,450

509

35.1%

0.98

Black women (intersectional)

890

231

26.0%

0.68

Native Hawaiian/Pacific Islander women

31

9

29.0%

0.76 (n=31 — unstable)

Read the table the way the auditor must. Single-axis ratios look tolerable — everything at or above 0.81. The intersectional row is where the story is: Black women at 0.68, a disparity invisible in both the sex table and the race table separately. LL144 made intersectional reporting mandatory precisely because Gender Shades showed failures concentrate at the intersections.

Then the small-n row: 31 applicants means one changed outcome moves the ratio by three points. The audit must publish the number and may annotate instability — what it must not do is quietly drop inconvenient categories. And remember the legal geometry from the state-laws module: LL144 sets no pass/fail line. The 0.68 does not violate LL144 — publishing it, with notice, satisfies LL144. What the 0.68 does is hand the EEOC and the plaintiffs’ bar a Title VII disparate-impact exhibit, flag four-fifths rule territory under the federal guidelines, and oblige you, under your own charter, to open a MANAGE-function response: feature analysis, threshold review, vendor escalation under your GOVERN 6 clauses.

Close out the audit with the two requirements teams forget. Independence: the auditor cannot be the vendor, and cannot be anyone with a stake in the tool’s continued use — an internal data scientist reporting to the product owner fails the test. Publication and notice: the summary of results, with the distribution date of the tool, goes on the careers site; the 10-business-day candidate notice machinery must be running before the first NYC requisition opens, with the alternative-process request path actually staffed.

Tool: Bias Audit Lab — Now run the audit yourself: same math, live data — tweak the screening threshold and watch the intersectional ratios and small-n instability respond.

Interactive checkpoint quiz (1 questions) — open this page in a browser to take it.