Frontier safety frameworks — and how to read them as a buyer

Lesson 5 of 5 in Third-Party AI, Generative AI, Agentic AI, and Frontier Governance.

At the top of the supply chain sit a handful of frontier labs whose models power much of what you procure. Their safety governance is therefore your upstream risk posture — and since roughly 2023 the labs have formalized it in public documents: Anthropic’s Responsible Scaling Policy (RSP), OpenAI’s Preparedness Framework, Google DeepMind’s Frontier Safety Framework. At the 2024 Seoul AI Safety Summit, sixteen companies committed to publishing frameworks of this kind.

The shared architecture is if-then scaling: define capability thresholds — measurable levels of dangerous capability, typically in CBRN uplift (could the model meaningfully help build biological or chemical weapons?), offensive cyber, and autonomous AI R&D (could the model dramatically accelerate its own improvement?) — then commit that if evaluations show a model crossing a threshold, then specified deployment safeguards and security standards must be in place before (or instead of) release. Anthropic’s version names the rungs: AI Safety Levels, modeled deliberately on biosafety levels. ASL-3, which Anthropic activated for its frontier models from 2025, pairs deployment standards (defenses against misuse of the dangerous capability) with security standards (protecting the model weights from theft — a stolen frontier model has no guardrails at all).

The honest framing: these are voluntary corporate policies, versioned and amendable by their authors. They are simultaneously the most substantive frontier-risk governance currently operating and a structure that would not survive a genuine race-to-the-bottom without regulatory backing. Both halves of that sentence are true, and a professional holds them together.

The three flagship frontier safety frameworks compared (verify current versions before relying on details)
FrameworkCore mechanismRisk domains trackedWhat crossing a threshold triggers

Anthropic — Responsible Scaling Policy (v3.4, July 2026)

AI Safety Levels (ASL) with capability thresholds; models must meet the safeguards of their level before deployment

CBRN uplift; autonomous AI R&D; with cyber and other domains tracked via evaluations

ASL-3+: hardened deployment standards (misuse defenses, defense-in-depth) and security standards (weight protection); semiannual published Risk Reports

OpenAI — Preparedness Framework (v2, Apr 2025)

Tracked risk categories scored against High / Critical capability thresholds

Biological & chemical; cybersecurity; AI self-improvement

High capability requires safeguards before deployment; Critical requires safeguards during development; internal safety advisory review

Google DeepMind — Frontier Safety Framework (v3, 2025)

Critical Capability Levels (CCLs) evaluated at compute/capability milestones

CBRN, cyber offense, harmful manipulation, ML acceleration/misalignment

Reaching a CCL triggers mitigation plans for deployment and security before further scaling or release

How should an enterprise buyer actually read these documents? Not as marketing, and not as scripture — as due-diligence exhibits answering four questions:

  1. Are the thresholds defined sharply enough to fail? A threshold no evaluation could ever trip is a press release. Look for named capability levels, described evaluation methods, and published results or risk reports.
  2. What happens when evaluations say "crossed"? The credible frameworks commit to specific deployment and security standards — and, critically, to pausing or restricting if safeguards are not ready. Check whether the framework has ever visibly bound the company’s behavior (delays, restricted releases, activated safeguards).
  3. Who verifies? Self-evaluation with no external red-teaming, no third-party assessment, and no publication commitment is the weakest configuration. Published risk reports and external testing arrangements strengthen it.
  4. What does it mean for your contract? A vendor’s frontier-lab supplier operating under a safety framework is an input to your own tiering — but it never substitutes for your deployment-level controls. The lab’s ASL-3 security standard does not ground your chatbot or gate your agent’s payments.

The statutory layer is arriving around the voluntary one. In the EU, the GPAI regime (taught in full in the transparency and GPAI module) has applied to new models since August 2025: Art 53 documentation and copyright-policy duties for all GPAI providers, Art 55 evaluations, adversarial testing, and incident reporting for systemic risk (GPAI) models, a voluntary Code of Practice (July 2025) as the compliance on-ramp, and a compliance deadline of 2 August 2027 for legacy models already on the market. What this means for you as a buyer: your frontier vendors now owe you, as a downstream provider, documentation under Art 53(1)(b) — ask for it by name.

In the US, California’s SB 53 (effective January 2026, per training knowledge — verify details) turned pieces of the voluntary architecture into law for large frontier developers: published safety frameworks, transparency reports, and critical-incident reporting, with whistleblower protections. The direction of travel is consistent: what labs volunteered in 2023–24 is becoming what statutes require.

Open-weight models flip the responsibility geometry. Download and self-host an open-weight model and there is no vendor operating the guardrails, no provider-side incident response, and no API to shut off — you absorb the duties a closed-model vendor would carry: safety evaluations for your use, abuse monitoring, update discipline when the base model is superseded, and license compliance (many "open" licenses carry use restrictions and redistribution terms). Fine-tune it and deploy it in a high-risk use, and the Art 25 role-switch from the procurement lesson applies with full force — except this time there is no upstream provider contract to lean on.

The frontier governance thread

  • 2025-02-01Frontier safety frameworks become table stakes:

    Following the Seoul commitments, major labs publish or update frontier safety policies (capability thresholds, evaluation gates, deployment mitigations) ahead of the Paris summit.

  • 2023-11-01Bletchley Park AI Safety Summit:

    28 countries + the EU — including the US and China — sign the Bletchley Declaration on frontier-AI risk. The summit series begins.

  • 2024-05-21Seoul AI Summit:

    Frontier labs sign safety commitments — publish risk frameworks or explain why not. The summit series turns from declarations to developer promises.

  • 2025-02-10Paris AI Action Summit:

    The series pivots from safety to action and investment; the US and UK decline to sign the final declaration — the divergence made visible.

  • 2026-02-19AI Impact Summit, New Delhi:

    The summit series lands in the Global South, centering development and inclusion; Geneva planned as the next stop (2027).

  • 2026-07-06First UN Global Dialogue on AI Governance:

    The Global Digital Compact’s forum convenes in Geneva — every state at one AI governance table for the first time.

Tool: AI Incident Tabletop — Stress-test this module’s controls: run an incident where your vendor’s genAI assistant fails at scale, against real regulatory clocks.

Key terms: responsible scaling policy, capability threshold, asl, frontier model, systemic risk (GPAI), open-weight model

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.