8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

Claude Opus 5: Full overview and first impressions

7/10
AI Eng.Ben BKJuly 25, 2026 at 11:07 AM16:58
Audio player
0:00 / 0:00

TL;DR

Opus 5 delivers near–top-tier AI performance at significantly lower cost, shifting the value frontier rather than raw capability leadership.

KEY POINTS

Performance gains across benchmarks

Opus 5 outperforms its predecessor and rivals on most of 13 major evaluations, including agentic coding, automation, and OS-level tasks. It shows one of the largest generational jumps in knowledge work and doubles Opus 4 in terminal-based coding. However, GPT 5.6 leads in at least one benchmark, highlighting that dominance is not absolute.

Cost-efficiency reshapes competition

The model’s biggest advantage lies in efficiency. On multiple benchmarks, Opus 5 achieves similar or better scores than Fable 5 at roughly half the cost. For example, high-end coding performance approaches parity while costing about $7–8 per task versus $20+ for competitors, effectively redefining the price-performance curve.

Mixed results in raw intelligence comparisons

Independent evaluations show Opus 5 closely matches GPT 5.6 in coding tasks, with no clear domination. This suggests progress is concentrated in optimization and usability rather than a leap in fundamental reasoning capability.

Breakthrough in interactive reasoning (ARC AGI 3)

On ARC AGI 3, a benchmark testing exploration and goal discovery in unknown environments, Opus 5 reaches 30.2%, a massive jump from near-zero levels earlier in 2026. However, this result required roughly $20,000 in compute, underscoring the importance of cost context in interpreting scores.

Scientific capabilities improve significantly

The model shows strong gains in life sciences, including +10.2 points in organic chemistry and +7.7 in protein-related tasks. Combined with high scores in biology benchmarks, it positions Opus 5 as one of the most capable publicly available models for scientific research.

Autonomous problem-solving behavior

Real-world tests highlight a pattern: the model builds missing tools itself. It has created custom computer vision pipelines, identified deep software bugs missed by others, and constructed validation systems when data was unavailable—demonstrating higher autonomy in complex workflows.

Safety-first design limits misuse

Opus 5 is engineered to detect vulnerabilities while limiting exploitation. It performs near top-tier in identifying security flaws but scores much lower in executing attacks. This reflects deliberate training choices to reduce harmful applications.

Reduced but persistent guardrails

Safety filters trigger 85% less often than in earlier models, improving usability. However, restrictions remain, particularly in cybersecurity contexts, where some users report blocked outputs even in legitimate testing scenarios.

Pricing and deployment

The model is priced at $5 per million input tokens and $25 per million output tokens, matching earlier versions but undercutting Fable 5, which costs roughly double. A faster mode offers 2.5× speed at higher cost.

User feedback highlights trade-offs

Early users praise its efficiency and long-task performance but note increased verbosity and occasional loss of focus. In some cases, the model prioritizes elaborate outputs over core task requirements, indicating room for improvement in instruction fidelity.

CONCLUSION

Opus 5 does not redefine the limits of AI intelligence but significantly lowers the cost of accessing high-level performance, intensifying competition on efficiency rather than raw capability.

Explain this
Full transcript

More from AI Eng.