Policy stacks, metrics, and making it stick
Lesson 5 of 5 in Designing the AI Governance Operating Model.
With structure and roles in place, the written layer follows a deliberate hierarchy — because a single document trying to be simultaneously inspirational, binding, and step-by-step ends up being none of the three. Each layer changes at a different speed and binds a different audience.
The policy architecture: five layers, five speeds
- Principles — Why — board-approved, changes rarely
A one-page statement of values and risk appetite: what the organisation will and will not do with AI. Approved at board level, stable for years. Everything below must trace to it.
- Policy — What — binding rules, reviewed annually
The enterprise AI policy and the acceptable-use policy. Mandatory statements: what requires review, who approves what, what is prohibited. Short enough that people actually read it — 5 to 10 pages, not 60.
- Standards — How well — measurable requirements per tier
Testable requirements the second line enforces: what a tier-1 validation must contain, which fairness metrics must be reported, retention periods for logs. Standards make "comply with the policy" auditable.
- Procedures — How, step by step — owned by teams
The operational runbooks: how to submit an intake form, how to run the monitoring review, the incident-response playbook. Change frequently; owned close to the work.
- Guidance — Help — non-binding, updated continuously
FAQs, worked examples, decision aids, office hours. Non-binding by design — the safety valve that keeps honest questions from becoming policy violations.
None of this operates in a vacuum. An AI governance program that duplicates what the enterprise already runs will lose the turf war and deserve to — the design move is integration, not invention. Enterprise risk management (ISO 31000) already owns the risk-appetite and register machinery: AI risks go into it as a category, not beside it. The privacy program already runs DPIAs under GDPR accountability: AI impact assessments share triggers and evidence with them rather than re-interviewing the same teams. The ISMS (ISO 27001) already certifies security controls: AI security requirements extend its statement of applicability. Banks route AI through existing model risk management; procurement adds AI due-diligence questions to existing vendor onboarding; ESG reporting picks up AI governance disclosures. If your organisation runs an ISO/IEC 42001 AI management system, that standard is this integration blueprint formalised — the AIMS module covers it in depth, and this module’s roles are the people who make its PDCA cycle turn. Likewise, everything here maps onto the NIST AI RMF Govern function — the RMF module shows the mapping.
Article 4 turns training from a nice-to-have into a compliance obligation — and role-based design is what makes it real rather than a mandatory e-learning click-through. Executives need enough to own risk appetite; committee members need to interrogate a model card; developers need the testing standards; frontline staff using AI outputs need to know the failure modes of the specific tools in their hands. One curriculum per audience, refreshed when the tools change, with completion tracked in the same registry that tracks the systems.
What you measure is what the board sees. A working metrics pack mixes three kinds of number. Coverage KPIs answer "how much of the estate is governed": % of AI systems inventoried, % assessed within SLA, % of high-tier systems with current validation. Effectiveness KRIs answer "is the governance real": committee rejection and conditions rates (a committee that approves 100% is a stamp), overdue-condition counts, drift alerts actioned within target, incident and near-miss counts. Efficiency metrics answer "is it sustainable": intake cycle time by tier, governance cost per use case. The same pack, aggregated, becomes the regulator-facing evidence trail — audit-readiness is a by-product of measuring honestly, not a separate project.
Resourcing follows the same proportionality logic as everything else. Early programs run on two to five dedicated people plus fractional legal, privacy, and security time; mature programs in regulated industries scale roughly with the count of high-tier systems, because those drive validation and monitoring load. Tooling has a predictable arc — every program starts on spreadsheets, and the migration trigger to a GRC platform or model-registry-integrated workflow is not size but synchronisation pain: when the spreadsheet no longer matches reality within a week of any change, buy or build the platform. Build-vs-buy turns on integration: buy the workflow shell, build the connectors into your own MLOps stack, because that is where the evidence lives.
And then there is the part no org chart fixes. Governance theater — structures that exist to be seen rather than to work — is the default failure state of this entire discipline, because every artifact in this module can be performed: a committee that meets and never rejects, a policy nobody can quote, training everyone clicks through, metrics defined so they cannot miss. The Dutch benefits scandal happened inside an administration with risk frameworks; Robodebt rolled on for years while frontline staff raised alarms that never travelled upward. The common thread is not missing structure — it is missing psychological safety: whether the engineer who suspects the training data is skewed, or the caseworker who sees the outputs are wrong, can say so and be heard without paying for it. Incentives close the loop — if delivery bonuses reward shipping and nothing rewards the person who stopped a bad launch, the organisation has told everyone what it actually wants.
Theater sign 1: the committee that never says no
Approval rates near 100% with thin minutes mean the real decisions happen elsewhere (or nowhere). Healthy programs show a visible rate of rejections and conditions — and can show conditions that were actually enforced, including at least one launch that was stopped or delayed.
Theater sign 2: metrics that cannot fail
"100% of submitted use cases reviewed" is a tautology — the denominator excludes everything that was never submitted. Honest coverage metrics use the discovered estate as the denominator, which requires the inventory work of the next module.
Theater sign 3: no bad news travels upward
Board packs showing only green for six consecutive quarters describe either a miracle or a filter. Ask when leadership last heard about an AI near-miss from inside the organisation rather than from a journalist. Robodebt’s Royal Commission found warnings existed for years — they just never survived the trip up the hierarchy.
Theater sign 4: governance headcount that only writes documents
If the governance team’s output is entirely policies, frameworks, and slide decks — and no one can point to a model that was changed, retested, or retired because of their challenge — the program is producing paper, not risk reduction.
Tool: AIMS Builder & Audit — Put this module to work: assemble a governance program piece by piece — roles, committee, policy stack, controls — and see where your design cracks under audit.
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.