Processes and gates: intake to production

Lesson 4 of 6 in The AI Center of Excellence: Organizing for AI Adoption.

Strip away the org charts and an AICoE is a pipeline with gates. Ideas enter at one end; governed, monitored production systems exit at the other; and at every gate the center asks the same two questions — is this still worth it? and is this still safe enough for its tier?

The pipeline starts with intake: one front door, one short form, no exceptions — including for executives, whose pet projects are the most common intake bypass. Microsoft’s CoE responsibilities list names this explicitly: a structured intake process with consistent criteria for business value, technical feasibility, and resource requirements, feeding a prioritized backlog. In practice the scoring trio is value × feasibility × risk: expected business value (quantified, owned by a named sponsor), technical feasibility (data exists, pattern exists, skills exist), and risk tier. That last one matters at intake, not later: the tier assigned on day one — using the same rubric the governance program defined — determines which gates apply, what evidence each gate demands, and who signs. Tiering late is the expensive mistake, because a tier-1 obligation discovered at the production gate can send a finished build back to the design stage. The mechanics of the form, the gating questions, and the registry schema get their own module (inventory and intake); here, what matters is that the CoE runs this machinery as one continuous flow.

The use-case pipeline: five stages, four gates

  1. Idea enters the front door

    A short intake form: problem, sponsor, data involved, people affected. Scored value x feasibility x risk; risk-tiered on the spot; logged in the registry from this moment — not from launch.

  2. Gate 1: worth a PoC?

    Portfolio decision against the prioritized backlog. Checks: quantified value hypothesis, named sponsor, data actually available, tier assigned. Most ideas should stop here — cheaply.

  3. Proof of concept

    Time-boxed (commonly 4-8 weeks), on the CoE platform with sandboxed data. Deliverable is evidence: does the approach work on our data at all?

  4. Gate 2: pilot with real users?

    Evaluation results against pre-agreed criteria, initial impact assessment for the tier, deployment cost estimate. The kill decision is a success: a PoC that dies here saved a pilot’s budget.

  5. Pilot

    Limited real users, real data, human oversight dialed high. Measures the value hypothesis — not just model metrics — plus incident drills and monitoring shakeout.

  6. Gate 3: production?

    The heavyweight gate: evaluation and fairness results vs release criteria, completed impact assessment, monitoring live, incident runbook tested, rollback plan, sign-offs per tier. High tiers go to the governance board; lower tiers use the delegated fast lane.

  7. Production and operation

    Monitoring against the thresholds set at gate 3; drift and incident alerts route to named owners; registry entry kept current.

  8. Gate 4: periodic re-review

    On cadence (annually, or per tier) and on triggers: major model change, new data source, incident, regulation change. Asks whether the system still earns its place.

  9. Retire or renew

    Systems that fail re-review are remediated or retired — data disposed per policy, registry entry closed. An estate that only ever grows is an estate nobody is reviewing.

Three pieces of machinery keep the pipeline honest. The AI system registry is the spine: every use case gets an entry at intake — owner, tier, status, data dependencies, gate history — and the entry lives as long as the system does. A registry populated only at launch is an obituary column; populated at intake, it is the single place where anyone (including a regulator) can see the whole estate and where every gate decision leaves its evidence.

Evaluation and release criteria are agreed before the build, per tier: which benchmarks, which fairness slices, which thresholds, what a red-team must fail to find. Criteria set after results exist get negotiated down to whatever the results happen to be — the pattern evaluation teams call teaching to the test, in reverse.

Exceptions and incidents are the pressure valves. A real pipeline needs a legitimate exception path — documented, time-boxed, owner-signed, risk-accepted at the right level, and logged in the registry — because the alternative to a governed exception is an ungoverned workaround. And when something breaks in production, escalation thresholds defined at gate 3 decide what the operating team handles versus what wakes the governance board: severity classes, notification clocks, and the kill-switch authority question answered in advance. The incident-response module drills this; the CoE’s job is making sure every system entered production with the runbook already written.

In a hub-and-spoke CoE, every pipeline responsibility has exactly one home: the hub (the central CoE), a spoke (the trained lead and team inside a business unit), or the platform (the shared infrastructure that enforces rules automatically, with the guardrails the cloud guardrails module covers). Microsoft’s AI-agents readiness guidance draws the same three-way split — a platform team managing the technical foundation and governance guardrails, workload teams owning their systems end to end, and the CoE as the advisory body driving strategy. Sorting responsibilities wrongly produces familiar pathologies: hub work pushed to spokes breeds inconsistency, spoke work hoarded by the hub rebuilds the bottleneck, and anything left to manual enforcement that the platform could automate will eventually be skipped under deadline pressure. Sort the deck:

Interactive sorting exercise: A hub-and-spoke AICoE is dividing responsibilities. Where does each one belong?

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.