The four functions, category by category
Lesson 3 of 5 in NIST AI RMF Deep Dive: Govern, Map, Measure, Manage.
The RMF Core is four functions — GOVERN, MAP, MEASURE, MANAGE — broken into 19 categories and 72 subcategories. The architecture matters more than the count: GOVERN is not a phase; it is the environment. It cross-cuts everything, building the organizational culture, roles, and policies inside which the other three run as a continuous cycle: MAP establishes context and surfaces risks, MEASURE analyzes and tracks them, MANAGE acts on them — and deployment feeds new context back into MAP. You never finish; you loop.
Before the functions, fix the terrain they operate on. The RMF describes the AI lifecycle in six stages, each with its own actors — and one activity, TEVV (test, evaluation, verification, and validation), that NIST deliberately refuses to confine to any single stage. Testing that happens once, at the end, is the failure pattern the whole framework is built against.
| Lifecycle stage | What happens | Who acts | Dominant RMF activity |
|---|---|---|---|
Plan & Design | System concept, objectives, context assumptions | Product owners, domain experts, governance leads | MAP 1 (context) and GOVERN — most cheap-to-fix risk decisions happen here |
Collect & Process Data | Sourcing, cleaning, labeling, documenting datasets | Data engineers, data scientists, privacy officers | MAP 4 (third-party data), MEASURE 2 (privacy, bias testing on data) |
Build & Use Model | Training, fine-tuning, model selection | ML engineers, modelers | MAP 2 (categorize the system), MEASURE 2 (validity checks in development) |
Verify & Validate | Pre-deployment TEVV against requirements and trustworthiness characteristics | TEVV specialists, red teams, auditors — ideally independent of builders | MEASURE at full intensity; MANAGE 1 go/no-go decisions |
Deploy & Use | Release into the real operating context | Deployers, integrators, end users, affected communities | MANAGE 1–2 (risk responses, benefit management), GOVERN 5 (feedback channels) |
Operate & Monitor | Ongoing performance tracking, drift detection, incident response, decommissioning | Operations teams, oversight humans, incident responders | MEASURE 3 (track risks over time), MANAGE 2.4/4 (decommissioning, continual improvement) |
GOVERN
The cross-cutting foundation: culture, accountability, and policy. GOVERN is where AI risk management becomes an organizational property instead of a heroic individual effort. Its six categories:
GOVERN 1 — Policies, processes, procedures, practices. The written machinery: policies that map legal and regulatory requirements onto the AI portfolio (1.1), that scale with risk (1.3 — this is where risk tolerance gets set and documented), and that mandate monitoring, decommissioning planning, and inventory. If a regulator asks “show me your AI policy”, GOVERN 1 is what they are asking for.
GOVERN 2 — Accountability structures. Named roles, empowered teams, trained people. Not “the data science team handles ethics” — documented lines showing who is responsible and accountable for each risk decision, with the authority and training to carry it.
GOVERN 3 — Workforce diversity and inclusion in decisions. Decision-making throughout the lifecycle draws on diverse perspectives, disciplines, and demographics — because homogeneous teams reliably miss the failure modes that hit people unlike them. (Note: the pending RMF revision is reported to target DEI-related text; the 1.0 category stands as written today.)
GOVERN 4 — Organizational culture. A safety-first mindset: teams that document by habit, communicate risks upward without career penalty, and treat critical challenge as a duty. NIST borrowed this deliberately from aviation and nuclear safety culture.
GOVERN 5 — Engagement with external stakeholders. Structured channels for feedback from users and affected communities — including appeal and recourse paths. This is the category state laws echo when they mandate consumer explanation and human-review rights.
GOVERN 6 — Third-party and supply-chain risk. Policies for risks arriving through vendors, pretrained models, and licensed data — plus contingency plans for third-party failures. When your foundation-model provider deprecates the API your product depends on, GOVERN 6 is why you had a plan.
MAP
Establish context; surface risks before you try to measure them. MAP is the anti-tunnel-vision function — the recognition that most AI failures are context failures: the model was fine, the deployment assumption was wrong. Its five categories:
MAP 1 — Context. Intended purposes, deployment settings, users, norms, expectations, and — crucially — documented assumptions and knowledge limits. A résumé screener built for engineering roles and quietly reused for warehouse staffing has violated MAP 1, whatever its accuracy.
MAP 2 — Categorization. What kind of system is this? The task, the methods, how humans will oversee and use outputs, where the system sits on the autonomy spectrum. Classification drives everything downstream — you cannot right-size controls for a system you have not characterized.
MAP 3 — Benefits, costs, and scope. What is the benefit, to whom, at what cost, and where does the appropriate application stop? Writing down the negative space — uses the system is not validated for — is the cheapest risk control in the framework.
MAP 4 — Third-party risks mapped. Risks and benefits of every third-party ingredient: pretrained models, licensed datasets, external APIs. You inherit the risks of components you did not build, and MAP 4 refuses to let “the vendor handles that” stand in for analysis.
MAP 5 — Impact characterization. Likelihood and magnitude of impacts on individuals, groups, communities, organizations, and society — the analytical bridge to MEASURE. This is where the impact-assessment documents that state statutes demand are born.
MEASURE
Analyze, assess, and track the risks MAP surfaced. MEASURE is where claims meet evidence. Its four categories:
MEASURE 1 — Methods and metrics. Choose appropriate quantitative and qualitative approaches — and, NIST insists, deal honestly with risks that resist quantification. Psychological harm from a companion chatbot has no clean metric; the RMF’s answer is structured qualitative assessment, not omission. What cannot be measured well must be documented as such, not dropped from the risk register.
MEASURE 2 — Evaluate against every trustworthiness characteristic. The workhorse category: its subcategories walk the seven characteristics one by one — validity and reliability, safety, security and resilience, transparency and accountability, explainability, privacy, and fairness with bias evaluated (including disaggregated, subgroup-level testing). This is the category a bias audit, a red-team exercise, and an accuracy benchmark all report into.
MEASURE 3 — Track risks over time. Mechanisms — dashboards, drift monitors, incident logs — that follow identified risks after deployment, including risks that were accepted rather than mitigated. Accepted risk without tracking is just forgotten risk.
MEASURE 4 — Validate the measurements themselves. Feedback from domain experts, users, and affected communities on whether the metrics actually capture what matters in context. A fairness metric can be computed perfectly and still measure the wrong thing — MEASURE 4 is the framework auditing its own ruler.
MANAGE
Allocate resources and act on what you mapped and measured. MANAGE closes the loop from analysis to decision. Its four categories:
MANAGE 1 — Prioritize and respond. For each significant risk, choose and document a response: mitigate (engineer the risk down), transfer (insurance, contractual allocation), avoid (do not deploy, or withdraw the feature), or accept (proceed, documented, within tolerance). The quartet comes straight from classical risk management — the discipline is writing the choice down and owning it.
MANAGE 2 — Maximize benefits, minimize harms. Resource the treatments, plan for sustained value — and maintain deactivation and decommissioning plans (MANAGE 2.4): mechanisms to supersede, disengage, or shut down a system that misbehaves, with minimal collateral damage. If turning the system off would be a crisis, that fact is itself a top-tier risk.
MANAGE 3 — Third-party risk, managed and monitored. The operational sequel to GOVERN 6 and MAP 4: pretrained models and vendor components get risk treatments and ongoing monitoring, not one-time onboarding checks.
MANAGE 4 — Document, respond, recover, improve. Post-deployment monitoring plans, incident response and recovery, communication plans for affected parties, and continual improvement feeding lessons back into MAP. The paper trail this category produces is exactly what a defense under Texas TRAIGA — or a regulator’s reasonableness inquiry — will ask to see.
Interactive sorting exercise: Name that function: drag each real-world governance activity to the RMF function it belongs to.
Key terms: TEVV, risk tolerance, residual risk, human oversight, model drift
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.