TEVV, vendor teeth, and the night the incident hits
Lesson 4 of 5 in Capstone: Building and Defending a US AI Compliance Program.
Week seven is testing week. For Windy, TEVV means red-teaming against the twelve risk categories of the NIST Generative AI Profile — for a candidate-facing HR chatbot, four dominate: confabulation (Windy inventing benefits terms a court could hold Windrose to, as Air Canada learned in Moffatt), harmful bias (steering candidates by inferred demographics), information security (prompt injection via resume text — a real attack channel, since Windy reads documents candidates control), and data privacy (leaking one worker’s pay data to another). Every red-team finding lands in the risk register with an owner and a MANAGE disposition: mitigate, accept with sign-off, or block launch.
Benchmarks are the trap here. A vendor eval sheet showing “95% on safety benchmarks” tells you almost nothing about your deployment — benchmarks measure the model in the lab, not the system in the wild wearing your system prompt, your retrieval corpus, and your users. Pre-deployment testing must exercise the assembled system, and then continuous monitoring takes over: drift metrics on AdvanceScore’s approval-rate distributions, sampled human review of Windy transcripts, quarterly re-runs of the SiftIQ disaggregated analysis. Write the decommissioning triggers now, in cold blood — the metric thresholds at which a system comes down — because no one makes that call well during an incident.
Vendor management is the same discipline pointed upstream (GOVERN 6, MAP 4). The documentation Colorado obliges SiftIQ’s developer to hand over from January 1, 2027 — intended uses, training-data categories, limitations, human-review instructions, update notices — is the floor; your contract should demand it now, plus what the statutes forgot: audit-grade data access for LL144, notification of model updates before they ship (a silently retrained SiftIQ invalidates your audit), TRAIGA-aligned representations about intended use, incident-notification flow-down on defined clocks, and indemnities sized to AG penalties. For AdvanceScore’s bank partner, SR 11-7 adds the vendor-model validation expectations: conceptual soundness review, ongoing monitoring, outcomes analysis — the bank will demand these of Windrose, so build them once and share.
Then, week nine, the phone rings. A worker in Sacramento reports that Windy, prompted cleverly, disclosed another worker’s wage-advance history and — worse — appears to have been doing so for days. Now the program you built either works under a clock or it does not.
Classify the incident, find the clocks
Interactive decision tree — outcomes:
- The playbook holds
Contain → preserve → classify → notify on every applicable clock (breach statutes, contractual flow-downs, the bank) → remediate → post-mortem feeding MAP. Evidence preservation is not optional housekeeping: the logs are what let you show the AG a contained, cured, documented incident instead of an unexplained one. This is MANAGE 4 plus GOVERN 4 under pressure.
- Wrong class, wrong playbook
Nothing here involves a protected-class outcome disparity. Misclassifying incidents sends you to the wrong clocks — discrimination findings trigger review of decisions and possible AG exposure; disclosure incidents trigger breach-notification law. Classification is the step that routes everything else, which is why the intake form forces it.
- Wrong regime, wrong reporter
SB 53’s incident duties attach to frontier developers for critical safety incidents in the catastrophic-risk sense. A deployer-level privacy leak is serious — but its clocks come from breach-notification and privacy law, not Cal OES. Knowing which regulator is NOT owed a filing is as much a part of competence as knowing which is; your duty to the frontier vendor runs through the contract.
- Dangerously wrong
There is no AI exemption from existing law — the theme of the entire sectoral module. Unauthorized disclosure of personal financial data triggers breach-notification statutes in every affected state, CCPA obligations, and contractual duties, regardless of the fact that a chatbot did it.
- The patch-and-pray anti-pattern
Quiet patching destroys the evidence that would have shown good faith, misses every notification clock, and converts a curable incident into concealment — the fact pattern regulators punish hardest and cure periods cannot save. Incidents are survivable; cover-ups are not.
Tool: AI Incident Tabletop — Run the full tabletop: an AI incident, live clocks, and every classification and notification decision on you.
Interactive checkpoint quiz (1 questions) — open this page in a browser to take it.