Anthropic launches Claude Opus 5 — state-of-the-art on Frontier-Bench and ARC-AGI 3, near-Fable-5 intelligence at half the cost, now default on Claude Max
Anthropic launched Claude Opus 5 on July 24, 2026 — the first Opus model in the fifth generation of Claude, priced at $5/$25 per million input/output tokens (same as Opus 4.8). Benchmark highlights: surpasses all models on Frontier-Bench v0.1 software engineering tasks at lower cost per task; scores 3× the next-best model on ARC-AGI 3 (novel problem solving); leads all models at any given cost on OSWorld 2.0 computer use…
Why people are talking about this Open context
Anthropic launched Claude Opus 5 on July 24, 2026 — the first Opus model in the fifth generation of Claude, priced at $5/$25 per million input/output tokens (same as Opus 4.8). Benchmark highlights: surpasses all models on Frontier-Bench v0.1 software engineering tasks at lower cost per task; scores 3× the next-best model on ARC-AGI 3 (novel problem solving); leads all models at any given cost on OSWorld 2.0 computer use (outperforming Fable 5 at one-third the cost); achieves ~1.5× the next-best model pass rate on Zapier AutomationBench end-to-end business task automation. Opus 5 is now the default model on Claude Max and the strongest on Claude Pro. Available on the Claude API as `claude-opus-5`, in Claude.ai, Claude Code, and Microsoft Foundry. Fast mode runs at 2.5× default speed for 2× the price. New beta capabilities alongside launch: mid-conversation tool changes (swap tools without invalidating prompt cache) and automatic fallbacks on the API (flagged classifier requests route to next-best model rather than blocking). Alignment: Anthropic's lowest automated behavioral audit score (2.3), below Opus 4.8, Sonnet 5, and Fable 5. Cybersecurity guardrails allow source-code vulnerability scanning but block binary scanning, penetration testing, and exploit generation; Cyber Verification Program members receive fewer restrictions. Note: ARC-AGI 3 cross-model comparisons are harness-dependent — OpenAI's retained-reasoning production configuration substantially raises GPT-5.6 Sol's score from generic-harness baselines.
Near-frontier intelligence at Opus cost and speed substantially expands the scope of autonomous enterprise agentic work; the combination of strong instruction-following, long-horizon persistence, and leading computer use at roughly half the cost of Fable 5 makes production multi-day agents economically viable for a much broader set of organizations.
'SOTA at half the price' claims rely on Anthropic-run benchmarks and early-access customer anecdotes; independent third-party evaluations under standardized harness configurations may reveal narrower advantages in specific domains; rapid model iteration (Opus 4.8 → Opus 5 within months) creates integration maintenance burden for teams on fast-following upgrade cycles.
Third-party independent benchmark evaluations of Opus 5 vs. GPT-5.6 Sol and Gemini 3.6 Ultra using standardized harness configurations; Automatic Fallbacks API adoption as an enterprise reliability pattern; Cyber Verification Program expansion to broader enterprise security teams.
- Run Opus 5 against target agentic coding and curation workloads to validate 'cost per successful task' claims vs. GPT-5.6 Terra
- Evaluate Automatic Fallbacks API feature (`claude-opus-5` → `claude-opus-4-8` on classifier flags) for production reliability in Claude-dependent workflows
- Track whether Opus 5 availability in Microsoft Foundry changes enterprise procurement dynamics vs. OpenAI Presence