Skip to main content

Claude Opus 5.5 sends most cyber tasks to Opus 4.8: what to test

Anthropic has tightened cyber routing in Opus 5.5. Teams evaluating security workflows should check which model actually handles their requests.

Two unmarked silicon chips separated by an upright glass panel on a dark surface
TechKili · AI-generated illustration with Cloudflare FLUX
Share this article:
In this article

Anthropic released Claude Opus 5.5 on September 22, 2026 with a consequential change for security work: the company says most cybersecurity tasks sent to the new model will be routed to Opus 4.8. Ordinary work to identify and fix bugs remains possible, but a team cannot assume that every result labeled as an Opus 5.5 session was produced by Opus 5.5. Anthropic's dated announcement establishes the release and the routing policy; The Verge reported it the same day.

That distinction matters to developers comparing model quality, cost and safety controls. The release is a new model, while the guardrail determines which model handles a sensitive request. A benchmark score for Opus 5.5 cannot, by itself, tell a security team what its own mix of tasks will receive.

What changed from Opus 5

Anthropic's Opus 5 announcement described cyber controls that allowed source-code vulnerability discovery while blocking some higher-risk work, including exploit generation. Flagged requests could already fall back to Opus 4.8. For Opus 5.5, Anthropic says it is applying a stronger class of cyber safeguards similar to those on Fable 5.1, and that most cybersecurity tasks will be rerouted to Opus 4.8. The difference is the expected reach of the intervention, not the invention of fallback itself.

Anthropic says its published benchmark runs used production safeguards; when these intervened, Opus 4.8 completed cyber tasks. The company also describes protections for biology and attempts to extract model capabilities. Its price list is $4 per million input tokens and $20 per million output tokens, below the Opus 5 list prices. The claim that typical workloads cost 40% less is based on Anthropic's tests. Teams should not apply that figure to a cyber workflow without measuring the actual route and token use.

Test the workflow, not only the model name

A useful evaluation starts with authorized examples of the team's real work: a routine code defect, a defensive investigation and any task that needs verified access. Record which requests are handled by Opus 5.5, which fall back, whether the answer remains useful, and the resulting time and cost. This is an editorial test plan, not a claim that Anthropic exposes identical routing details in every product. Keep any security testing within systems the team is permitted to assess.

Anthropic says it will expand its Cyber Verification Program to Opus 5.5 in the coming weeks. That is a future access path for vetted practitioners, not general availability of unrestricted cyber use today. A team that depends on those capabilities should check its present access before migrating a production workflow.

What the safeguards do not establish

The launch page describes internal behavioral audits and external pre-release evaluators, but the headline comparisons are predominantly Anthropic's own results. They do not measure whether a particular organization's security work improves after deployment. A stronger classifier also does not substitute for permissions, network isolation or human review. Our earlier report on Claude reaching real systems during cyber evaluations explains an infrastructure boundary failure; that incident should not be presented as proof of how Opus 5.5 behaves in normal use.

The immediate decision is narrow. For general coding, Opus 5.5 may be worth testing against the previous model on the team's own tasks. For security work, first establish the effective model route and access tier. Anthropic has announced a deployment policy and prices; independent results for each team's workload remain to be measured.

Sources