Running the review — and mapping it to everything else

Lesson 5 of 5 in The Generative AI Lens: AWS Well-Architected for GenAI.

Knowing the questions is half the job; running the review well is the other half. Three practices separate reviews that change workloads from reviews that produce a forgotten PDF.

Prepare evidence, not opinions. Before the session, gather what the questions will ask about: the latest GENOPS01 evaluation report, the CloudWatch dashboards, the guardrail configurations and their intervention metrics, the agent execution-role policies, the cost breakdown by model. A best practice gets ticked when someone can show it, not when the room feels good about it. The evidence-gathering itself is diagnostic — a control nobody can produce evidence for in an afternoon is a control that will not survive an auditor either.

Prioritise High-risk gaps, ruthlessly. The tool will hand you a mixed list of HRIs and MRIs. Resist the temptation to clear ten easy MRIs and declare progress — the lens authors marked GENOPS01-BP01 or GENSEC05-BP01 High because their absence is how workloads end up in incident reports. Every HRI gets an owner, a date, and a follow-up milestone.

Put the re-review on a trigger, not just a calendar. A scheduled cadence (commonly every six to twelve months) is the floor. The real triggers are events: a new model version behind your application, a new customisation, a new data source in the RAG corpus, a new agent capability. Each of those silently invalidates old answers — the milestone history in the WA Tool is how you prove to yourself, and later to others, that answers were refreshed when the workload changed.

Representative lens best practices, crosswalked (Annex A families per ISO/IEC 42001; verify exact control numbers against the current standard)
Lens best practiceISO/IEC 42001NIST AI RMFGuardrail Catalog entry

GENOPS01-BP01 — periodic stratified evaluation

Clause 9.1 monitoring & measurement; A.6 life-cycle verification and validation

MEASURE — valid and reliable behaviour, quantified in operation

Evaluation harness / grounding-check entries — the mechanism behind declared accuracy

GENSEC02-BP01 — implement guardrails on responses

A.6 design criteria; A.9 responsible use

MANAGE — treatment of identified output-harm risks; safe characteristic

Content filter, denied-topic, and grounding-check entries

GENSEC03-BP01 — control-plane and data-access monitoring

A.6.2.8 event-log recording

MEASURE + GOVERN — the evidence base and its accountability

Invocation-logging entry — same logs, third framework

GENSEC05-BP01 — least privilege for agentic workflows

A.9 responsible use; A.4 resources and roles

GOVERN + MANAGE — accountable oversight of autonomy

Agent permission-scoping and human-confirmation entries

GENOPS02 — operational-health monitoring

Clause 9.1; A.6 operation and monitoring

MEASURE — tracking performance in deployment

Monitoring and alert-threshold entries

GENSUS01 — minimise computational footprint

No direct control — nearest hook is the impact-assessment theme (A.5), where environmental effects can be assessed

GOVERN — organisational values expressed as policy; RMF names environmental impact among AI risks

No runtime entry — an architecture decision, not a request-path control

Two honest observations about that table. First, mappings are lossy: GENSUS has no clean 42001 counterpart, and 42001’s organisational clauses (leadership, competence, internal audit) have no lens counterpart — a gap in one framework is not a gap in the other, and pretending the tables line up perfectly is how crosswalks lose their readers’ trust. Second, the mapping runs through evidence: the same GENOPS01 evaluation report is a ticked best practice in the lens review, a clause 9.1 monitoring record in the AIMS, and a MEASURE artifact in an RMF profile. Write each artifact once, cite it three times.

So when do you reach for the lens, and when for ISO/IEC 42001? Wrong question — they are different instruments at different altitudes. The lens reviews one workload on one cloud: is this system well built, by AWS engineering standards, right now? The AIMS governs the organisation: are there policies, roles, impact assessments, audits, and improvement loops covering every AI system you run, certifiably? A company with brilliant lens reviews and no management system has excellent workloads and no governance; a certified AIMS whose workloads never face a technical review has governance on paper and unknown engineering underneath.

Run them as a loop: the AIMS decides which workloads exist and what risk appetite they operate under; the lens review pressure-tests each workload’s engineering; unmet HRIs flow back into the AIMS as nonconformities or improvement actions (clause 10), and lens milestones become evidence the internal audit (clause 9.2) can sample. One system manages, the other measures — and both quote the same dashboards.

Key terms: AI management system, internal audit, nonconformity (major / minor), crosswalk

Tool: GenAI Lens Index — Browse the full best-practice index: every GENOPS-to-GENSUS question and best practice with a plain-language summary and a link to the official page — the reference you will want open during a real review.

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.