Key Takeaways
- Anthropic says Claude Opus 5 is available now and priced the same as Opus 4.8.
- The company positions it as a stronger model for coding, knowledge work, and long-running agentic tasks.
- Anthropic also says Opus 5 remains behind Mythos 5 on cybersecurity exploitation tasks.
What happened
Anthropic has introduced Claude Opus 5 and says the model is available today across its platforms. The company describes it as a more thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
In Anthropic’s reporting, Opus 5 is the new state-of-the-art on coding and knowledge-work evaluations such as Frontier-Bench and GDPval-AA, while still trailing Mythos 5 on cybersecurity tasks. Anthropic also says the model is designed for everyday use, works more efficiently than other models, and is now the default model on Claude Max and the strongest model on Claude Pro.
The launch centers on performance and cost-effectiveness. Anthropic says Claude Opus 5 improves meaningfully over Opus 4.8 at the same cost. The company highlights several benchmark results. On Frontier-Bench v0.1, it says Opus 5 surpasses all other models and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, Anthropic says the model gets within 0.5% of Fable 5’s peak score at half the cost per task.
The company also points to results beyond coding. On ARC-AGI 3, Anthropic says Opus 5 scored three times as high as the next-best model. On Zapier AutomationBench, it says the model’s pass rate was around 1.5 times the next-best model for the same cost per task, and that even at the lowest effort setting it passed more tasks than any other model. On OSWorld 2.0, Anthropic says Opus 5 outperformed every other model at any given cost and exceeded Fable 5’s best result at just over a third of the cost.
Anthropic says the model also improves on scientific research workflows, especially in life sciences, including structural biology, organic chemistry, and bioinformatics. The company says its gains were most notable on organic chemistry and protein-related tasks.
A second theme in the launch is agentic behavior. Anthropic says Opus 5 is stronger at verifying its own work and iterating until it succeeds. The company describes examples in which the model wrote its own computer-vision pipeline, found the root cause of a bug in an open-source package manager, and built a test harness when it lacked a live feed to validate code against.
Anthropic says early-access customers also saw improvements in difficult debugging, code review, financial research, legal review, presentation work, and monitoring-agent workflows. Across those examples, the company emphasizes consistency, reduced variance, and stronger performance on long-horizon tasks.
Why it matters
The main significance of Claude Opus 5 is not just that Anthropic has released another flagship model, but that it is framing the model around practical productivity and cost. The company repeatedly emphasizes better performance at the same or lower cost, which is a key metric for developers and enterprise customers deciding where to deploy AI in production.

That is especially clear in the software engineering results. Anthropic positions Opus 5 as its strongest model for coding and related agentic work, and says it is now the default on Claude Max and the top option on Claude Pro. If those claims hold up in real deployments, the model could matter most for teams looking for reliable multi-step coding assistance rather than purely conversational output.
The launch also shows Anthropic leaning into an important distinction in AI capability: higher general capability does not automatically mean higher risk in every domain. Anthropic says Opus 5 does not advance the frontier in risky dual-use capabilities and remains behind Mythos 5 in both biology research and offensive cybersecurity. At the same time, it says the model is its most aligned to date and its safest model yet in terms of avoiding reckless actions.
That safety framing matters because Anthropic is also shipping narrower guardrails and fallback behavior. The company says Opus 5’s cyber classifiers allow vulnerability discovery in source code but block binary-based scanning, penetration testing, and exploit generation. In Claude.ai, Claude Code, and Claude Cowork, flagged requests fall back to Opus 4.8 by default, and API users can enable automatic fallbacks.
For developers, the product changes may be as important as the benchmark claims. Anthropic is adding mid-conversation tool changes on the Claude Platform and automatic fallbacks on the API. Those updates suggest the company is trying to make the model easier to use in real workflows, especially where tool access or safety filters would otherwise interrupt a session.
What to watch
The biggest question is how Claude Opus 5 performs outside Anthropic’s own evaluations and customer quotes. The source material is rich in benchmark claims, but it does not include independent testing, so real-world performance remains to be seen.
It will also be important to see how widely developers adopt the model for production coding, debugging, and agentic automation. Anthropic says Opus 5 is especially strong on long-running, multi-step work, but that is exactly the kind of capability that tends to be judged by reliability over time rather than a single benchmark run.
Another area to watch is how the safety and fallback setup affects availability. Anthropic says some cybersecurity and biology-related requests will route differently depending on the product, the classifier, and whether a user is in the Cyber Verification Program. That means the practical experience of using Opus 5 may differ depending on the workflow.
Finally, pricing and speed are part of the story. Anthropic says Opus 5 is priced the same as Opus 4.8, and that Fast mode runs around 2.5 times the default speed at twice the base price. How customers balance those options will likely determine whether Opus 5 is seen as a routine upgrade or a specialized premium model.



