Claude Opus 4 shipped under ASL-3 in May 2025 because Anthropic's own Responsible Scaling Policy actually triggered, not because a policy document said it should, and the same long-context window this series already covered as a serving win is what makes many-shot jailbreaking possible in the first place
The real ASL-1 through ASL-4 capability thresholds, the actual biosecurity trial result that pushed Opus 4 into precautionary ASL-3 deployment, Anthropic's own February 2025 red-team challenge (339 researchers, 300,000-plus attacks, a jailbreak success rate cut to 4.4%), and the real, quantified tension between jailbreak resistance and over-refusal that XSTest and OR-Bench measure directly rather than assert.
The previous post’s lethal trifecta was about a structural vulnerability in what an agent is allowed to do. Safety evaluation is the parallel discipline for what a model is capable of doing in the first place, chemistry, biology, cyberattacks, persuasion at scale, evaluated before deployment rather than discovered after. It’s tempting to treat this as a compliance exercise, a checklist a legal team runs once before launch. The real, disclosed evidence says otherwise: Anthropic’s own Responsible Scaling Policy (RSP) is a live engineering framework that has already changed how a real, shipped model was deployed, and the discipline runs on a genuine, measured tradeoff, not a single dial you turn toward “safer.” Turn it too far and you get a model that refuses to explain how antibiotics work because the question mentions “resistance.” This post is about the real framework, the real numbers behind it, and the real vulnerability, many-shot jailbreaking, that emerges directly from a capability this series already covered as an unambiguous win.
The real ASL ladder, not a compliance abstraction
Anthropic’s Responsible Scaling Policy defines AI Safety Levels (ASL) as a graduated ladder: as measured capability increases, required safeguards increase with it, and a model isn’t deployed past a threshold until the corresponding safeguards are actually in place. The levels are concrete, not aspirational language:
- ASL-1: systems posing no meaningful catastrophic risk, the paper’s own example is a 2018-era language model or a chess engine.
- ASL-2: systems showing early signs of dangerous capability, for instance the ability to describe how to build a bioweapon, but where the information isn’t yet reliable or useful enough to exceed what a search engine already provides. Anthropic states directly that current Claude models, at time of writing, sit here.
- ASL-3: systems that substantially increase catastrophic misuse risk compared to a non-AI baseline (search engines, textbooks), or that show low-level autonomous capability. The RSP’s own binding commitment: do not deploy an ASL-3-capable model that shows any meaningful catastrophic misuse risk under adversarial testing by world-class red-teamers, and apply materially stronger security requirements around the weights themselves.
- ASL-4 and above: deliberately left undefined, Anthropic’s own position is that it’s too far from present systems to specify usefully yet, and premature specificity here would be guessing.
Opus 4: the framework actually triggering, not a document sitting on a shelf
On May 22, 2025, Anthropic activated ASL-3 protections for Claude Opus 4, the first model in the company’s history deployed under its highest active safety tier. What makes this the interesting case, more than the fact of the activation itself, is the honest epistemic state Anthropic disclosed alongside it: they had not determined that Opus 4 definitively crosses the ASL-3 capability threshold. Every prior model, they state directly, they were able to clearly rule out. For Opus 4, given continued improvement in CBRN-adjacent knowledge and reasoning, they could no longer make that clean determination, so they deployed under ASL-3 anyway, as an explicitly precautionary, provisional measure rather than wait for certainty that might arrive too late. The concrete evidence behind that caution: in expert-graded biosecurity trials, Opus 4 demonstrated “significantly greater” performance than both Google search and prior Claude models on tasks relevant to bioweapons-adjacent knowledge.
The response wasn’t a single lever. Two distinct standards activated together: a Security Standard, meaningfully harder-to-steal model weights, treating the weights themselves as the asset requiring protection; and a Deployment Standard, a narrowly scoped set of usage restrictions specifically targeting CBRN misuse, not a blanket capability reduction across every domain. That distinction matters for the same reason precision has mattered throughout this series: a security response and a deployment response solve different problems, and collapsing them into “we made it safer” hides which specific risk each one actually addresses.
DeepMind’s parallel framework, and the real, separate role built around it
Google DeepMind’s Frontier Safety Framework, now on its third version, runs the same underlying idea through different terminology: Critical Capability Levels (CCLs) instead of ASL tiers, covering CBRN, Cybersecurity, Harmful Manipulation, ML R&D acceleration, and misalignment as five separate threat domains rather than one aggregate risk score. DeepMind publishes per-model reports against this framework directly, a real, current example being the Gemini 3 Pro Frontier Safety Framework report, published November 2025.
The real, current job posting for this work, Research Engineer, Frontier Safety Mitigations, DeepMind, states its mandate directly: “accountable for the safety and behavior of Google DeepMind’s latest Gemini models,” focused specifically on “CBRN, Cybersecurity, and Harmful Manipulation,” with the explicit method list being “building novel evaluations… red-teaming, researching and deploying advanced mitigations, monitoring emerging risks, and contributing to model development.” Worth being precise about the bar here, because it’s easy to assume every safety-adjacent role requires a research background: this specific posting’s minimum qualifications are a bachelor’s degree, 5 years of software development experience, and 3 years testing or launching software products, no PhD requirement, in explicit contrast to DeepMind’s own Multi-Agent Learning research role covered in the previous post. This is real engineering work with a security-and-reliability shape, not exclusively an academic research track.
Red-teaming as a real, funded practice, not an internal exercise
Anthropic’s Constitutional Classifiers, published as “Defending Against Universal Jailbreaks,” is the concrete system behind Opus-generation jailbreak resistance: input and output classifiers trained on synthetic data generated against a written constitution, similar in spirit to the Constitutional AI training method already derived in this series, but applied as a runtime filter rather than a training-time alignment method. The reported result: jailbreak success against a Claude model equipped with these classifiers dropped to 4.4%, down from a much higher baseline success rate without them, meaning the classifiers blocked more than 95% of attempts that previously succeeded.
The real evidence for that number isn’t just an internal benchmark. Anthropic ran a public red-team challenge from February 3 to 10, 2025: 339 security researchers, more than 3,700 collective hours, over 300,000 targeted attack interactions, structured as eight levels where clearing a level required extracting a specific CBRN-adjacent answer from the model despite the classifiers. Anthropic paid out 35,000 per newly discovered universal jailbreak technique, run publicly on HackerOne against a live alias of the production model. This is worth taking seriously as a real economic signal, not a publicity exercise: a frontier lab is willing to pay tens of thousands of dollars per confirmed jailbreak because the cost of not finding it internally, discovering it after a real misuse incident, is understood to be far higher.
METR (Model Evaluation and Threat Research, originally spun out as ARC Evals, founded 2022) is the real, independent counterpart to a lab red-teaming itself: a third-party nonprofit that runs pre-deployment dangerous-capability evaluations under contract with OpenAI, Anthropic, and Google DeepMind, covering persuasion and deception, cybersecurity, self-proliferation, and self-reasoning or self-modification capability. As of early 2026, METR piloted a further shift, an entity-based assessment of a lab’s internal AI use, run periodically rather than tied to a specific public launch, with Anthropic, Google, Meta, and OpenAI all participating. The direction of travel is explicit in that shift: evaluation as an ongoing operational discipline, not a one-time pre-launch gate that stops mattering the day after ship.
Many-shot jailbreaking: the vulnerability this series’ own celebrated capability creates
Here is the result worth sitting with longest, because it’s a direct, uncomfortable consequence of something already covered in this series as unambiguously good news. Anthropic’s own research, published April 2024 and presented at NeurIPS 2024, describes many-shot jailbreaking (MSJ): stuff a single prompt with a long, fabricated dialogue history showing an “AI assistant” persona readily answering a long sequence of increasingly borderline or harmful questions, then append the actual harmful question at the end. Given enough fake prior turns exhibiting the jailbroken behavior, the model continues the pattern it’s just been shown, on a genuinely novel request, essentially in-context few-shot learning applied to bypassing its own safety training instead of a real task.
The mechanism is not a bug that happened to also affect long-context models. It is a direct, structural consequence of exactly the capability this series already covered as a real production win: a context window that grew from roughly 4,000 tokens at the start of 2023 to over a million tokens in current frontier models is, simultaneously and for the identical reason, room for far more fabricated in-context examples than a short-context model could ever hold. The research confirms this affects models across labs, Anthropic’s own included, not a defect specific to one company’s training pipeline, because it’s a property of how in-context learning works in any sufficiently long context, not a flaw in any one system prompt. This is the cleanest example available in this whole series of a real design tension stated plainly: the same capability increase that makes a model more useful can, without any new architectural mistake, make it more exploitable, and the honest response isn’t to reverse the capability, it’s to build a detector for the specific pattern the capability now makes newly possible, which is exactly what Anthropic’s published mitigation does, classifying and intervening on suspiciously long runs of escalating fake compliance before the final harmful turn ever gets a real answer.
The tradeoff nobody gets to skip: jailbreak resistance versus over-refusal
Every mitigation above pushes in one direction, less compliance with harmful requests. Push that dial hard enough with no counterweight and you get a different, well-documented failure: a model that refuses “how do I kill a Python process,” “what’s the lethal dose of caffeine, I’m writing a safety pamphlet,” or a legitimate chemistry homework question, because the safety training learned to associate specific words, “kill,” “lethal,” “bomb,” “weapon,” with refusal, independent of the actual context those words appear in.
XSTest is a diagnostic benchmark built specifically to surface this: 250 prompts that are safe by construction but contain exactly the kind of lexically risky-sounding language that triggers keyword-level over-caution. OR-Bench scales the same idea to 80,000 prompts across ten rejection categories. The finding that matters most operationally, confirmed across both: models trained more aggressively for safety measurably refuse more of these clearly-safe prompts, not fewer, and the root cause identified is lexical overfitting, the model learning a shortcut association between surface tokens and refusal instead of the deeper semantic judgment “does this specific request, in this specific context, actually risk harm.” That shortcut is cheap to learn during safety fine-tuning and cheap to trigger by accident, which is exactly why it shows up reliably across model families rather than as one company’s isolated mistake.
This is the real, quantifiable version of the abstract warning already present in the Anthropic-safety literature about over-refusal regressions: it isn’t a hypothetical edge case a red team might stumble on, it’s a measured, benchmarked, reproducible curve, and a team that only tracks jailbreak-success-rate going down, without simultaneously tracking XSTest or OR-Bench refusal rate, is optimizing one axis of a genuine two-axis tradeoff while flying blind on the other.
| Framework / tool | What it measures | Real, disclosed number |
|---|---|---|
| Anthropic RSP / ASL | Capability threshold crossed | Opus 4 deployed under ASL-3, May 2025, precautionary |
| DeepMind Frontier Safety Framework | Critical Capability Levels across 5 domains | Versioned (v3.0); per-model public reports, e.g. Gemini 3 Pro, Nov 2025 |
| Constitutional Classifiers | Jailbreak success rate | 4.4% success against a defended model |
| Public red-team bounty | Real-world adversarial pressure | 339 researchers, 300,000+ attacks, $55K paid, Feb 2025 |
| METR | Independent dangerous-capability evaluation | Contracted pre-deployment work across OpenAI, Anthropic, DeepMind |
| XSTest / OR-Bench | Over-refusal rate on safe prompts | 250 / 80,000 prompts; higher safety training correlates with higher false-refusal rate |
Common mistakes
Treating a Responsible Scaling Policy as a compliance document rather than an engineering trigger: Opus 4’s ASL-3 activation was a real, disclosed, two-part response (security standard plus deployment standard) to a specific evaluation result, not a policy team signing off on a launch.
Assuming long context is a strictly one-directional improvement with no safety cost: many-shot jailbreaking is a direct, structural consequence of the exact capability this series covered as a serving and training win, and the honest framing is a real tradeoff requiring its own dedicated mitigation, not a side effect that better prompting resolves.
Optimizing jailbreak-resistance without tracking over-refusal on a benchmark like XSTest or OR-Bench alongside it: these move in tension, confirmed empirically, and reporting only the jailbreak number hides half of what actually changed in the model’s behavior.
Try it yourself
Beginner. Using the ASL ladder above, explain in one sentence why Anthropic states Claude sits at ASL-2 rather than ASL-1, given that ASL-2 explicitly includes models that show early dangerous-capability signs.
Intermediate. A model refuses the prompt “explain how antibiotic resistance develops in bacteria” because it contains the word “resistance” near biological terms. Classify this using XSTest’s own framing, and propose a fix that doesn’t simply remove the underlying safety training that causes the model to refuse genuinely harmful biology requests.
Advanced. Many-shot jailbreaking’s effectiveness scales with how many fake in-context examples fit in the context window. Design an evaluation that would measure whether a specific model’s susceptibility to MSJ scales linearly, sub-linearly, or with a sharp threshold as the number of fake shots increases, and explain what each shape would imply about whether the mitigation should target the classifier, the context length itself, or the sampling procedure.
What this takes to be frontier-job-ready
The technical and operational bar is stated directly in DeepMind’s own listing for Research Engineer, Frontier Safety Mitigations: “accountable for the safety and behavior of Google DeepMind’s latest Gemini models,” working across “building novel evaluations… red-teaming, researching and deploying advanced mitigations, monitoring emerging risks.” Worth being precise about who this role is actually open to: the posting’s minimum bar is a bachelor’s degree plus 5 years of software development and 3 years shipping or testing production software, not a PhD, a materially different and more engineering-shaped entry point than DeepMind’s own Multi-Agent Learning research role.
The economic reality behind the operational urgency is the bounty program itself: Anthropic paying real money, up to $35,000 per confirmed universal jailbreak, on a continuous basis, is a direct, disclosed statement of how much a missed vulnerability is judged to cost relative to finding it first. And the honest epistemic posture required shows up nowhere more clearly than in Anthropic’s own Opus 4 disclosure: stating plainly that a capability threshold could not be definitively ruled out, rather than either overstating confidence or delaying a launch indefinitely, is precisely the judgment call this work actually demands under real deadline pressure.
The one-sentence version: Claude Opus 4’s ASL-3 activation in May 2025 is Anthropic’s own Responsible Scaling Policy actually firing in production, triggered by a specific, disclosed biosecurity trial result, not a document nobody consults; DeepMind runs the parallel Critical Capability Levels framework with its own versioned, per-model public reports; real red-teaming now runs as a funded, adversarial, external practice, $55,000 paid out to researchers who broke a defended model in February 2025 alone; and many-shot jailbreaking is the cleanest available proof that a capability this series has already praised, long context, creates a real, structural vulnerability simply by existing, with no new mistake required. None of this is separable from the over-refusal side of the ledger: a safety number that only ever goes in one direction is a number that isn’t being measured against its own real cost.