
Tech • AI • Robotics
Opus 5 delivers near–top-tier AI performance at significantly lower cost, shifting the value frontier rather than raw capability leadership.
Opus 5 outperforms its predecessor and rivals on most of 13 major evaluations, including agentic coding, automation, and OS-level tasks. It shows one of the largest generational jumps in knowledge work and doubles Opus 4 in terminal-based coding. However, GPT 5.6 leads in at least one benchmark, highlighting that dominance is not absolute.
The model’s biggest advantage lies in efficiency. On multiple benchmarks, Opus 5 achieves similar or better scores than Fable 5 at roughly half the cost. For example, high-end coding performance approaches parity while costing about $7–8 per task versus $20+ for competitors, effectively redefining the price-performance curve.
Independent evaluations show Opus 5 closely matches GPT 5.6 in coding tasks, with no clear domination. This suggests progress is concentrated in optimization and usability rather than a leap in fundamental reasoning capability.
On ARC AGI 3, a benchmark testing exploration and goal discovery in unknown environments, Opus 5 reaches 30.2%, a massive jump from near-zero levels earlier in 2026. However, this result required roughly $20,000 in compute, underscoring the importance of cost context in interpreting scores.
The model shows strong gains in life sciences, including +10.2 points in organic chemistry and +7.7 in protein-related tasks. Combined with high scores in biology benchmarks, it positions Opus 5 as one of the most capable publicly available models for scientific research.
Real-world tests highlight a pattern: the model builds missing tools itself. It has created custom computer vision pipelines, identified deep software bugs missed by others, and constructed validation systems when data was unavailable—demonstrating higher autonomy in complex workflows.
Opus 5 is engineered to detect vulnerabilities while limiting exploitation. It performs near top-tier in identifying security flaws but scores much lower in executing attacks. This reflects deliberate training choices to reduce harmful applications.
Safety filters trigger 85% less often than in earlier models, improving usability. However, restrictions remain, particularly in cybersecurity contexts, where some users report blocked outputs even in legitimate testing scenarios.
The model is priced at $5 per million input tokens and $25 per million output tokens, matching earlier versions but undercutting Fable 5, which costs roughly double. A faster mode offers 2.5× speed at higher cost.
Early users praise its efficiency and long-task performance but note increased verbosity and occasional loss of focus. In some cases, the model prioritizes elaborate outputs over core task requirements, indicating room for improvement in instruction fidelity.
Opus 5 does not redefine the limits of AI intelligence but significantly lowers the cost of accessing high-level performance, intensifying competition on efficiency rather than raw capability.
Explain this