Anthropic released Claude Opus 5 on July 24, 2026. It immediately took the top spot on Artificial Analysis's Intelligence Index at a score of 61, and led the Agentic Index at 55.3 — the two most-watched independent benchmark suites for frontier models right now.
What Shipped
Beyond benchmark rankings, Opus 5 introduces critical platform features designed for sustained agentic work:
Extended Context & Output: Features a standard 1M-token context window with a 128k maximum output token capability.
Adaptive Reasoning & Mid-Conversation Tool Updates: Adapts thinking depth automatically on each turn while allowing developers to add or swap API tools dynamically mid-session without invalidating prompt caches.
Self-Verification Loops: Built-in self-correction capabilities reduce hallucination and eliminate the need for manual verifier subagent prompts.
The Price Angle
Opus 5 is priced at $5 per million input tokens and $25 per million output tokens. That is half the price of Anthropic's Mythos-tier Claude Fable 5.
For enterprise engineering and product teams, this price adjustment changes the math: top-tier intelligence ratings are no longer coupled to premium API pricing. Lower prompt cache minimums (512 tokens down from 1,024) further reduce running expenses for high-frequency workflows.
Where It Sits in the Field
Opus 5 lands during a crowded month for frontier model drops:
GPT-5.6: OpenAI's Sol/Terra/Luna tier family staged rollout (starting July 9).
Grok 4.5: xAI's coding-focused release (July 8), now integrated under the unified SpaceXAI brand.
Currently, Opus 5 leads both models on Artificial Analysis rankings, outperforming competing labs on multi-file software refactoring (SWE-bench Verified) and complex autonomous agent execution (OSWorld 2.0).
What It Means for You
If your infrastructure already relies on Claude for software development or agentic task routing, Opus 5 provides an immediate upgrade at a lower per-token operational cost.
If you are evaluating providers from scratch, treat Opus 5 as the new baseline benchmark. Because independent benchmark rankings do not always mirror specialized real-world performance, run task-specific evaluations against your own codebase before shifting production traffic.