Claude Opus 5.5: 40% Cheaper Than Opus 5, 66.4% on Terminal-Bench 4.0
Anthropic released Claude Opus 5.5, the first model in the new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Input and output tokens are priced at $4 and $20 per million respectively — 20% less than Opus 5 — while cache reads, which Anthropic says make up most agentic and coding work costs, drop to $0.20 per million, a 60% reduction. Output generates more than 30% faster than Opus 5.
Benchmarks
Opus 5.5 leads across agentic coding, computer use, and knowledge work:
- Terminal-Bench 4.0: 66.4% (vs 55.8% Fable 5.1, 52.3% Opus 5, 57.9% GPT-6 Astra, 37.3% GPT-5.6 Sol)
- FrontierCode v1.1 (Main): 54.4% (Fable 5.1: 50.3%, Opus 5: 48.0%)
- CursorBench 4.0: 57.8% (Fable 5.1: 51.8%, Opus 5: 46.6%)
- GDPval-AA v2.1: 1846 (Fable 5.1: 1735, Opus 5: 1708)
- AutomationBench: 40.0% (Fable 5.1: 31.4%, Opus 5: 26.9%)
- Humanity's Last Exam: 67.7% with tools (Fable 5.1: 65.6%, Opus 5: 63.6%)
- Terminal-Bench-Science 0.1: 58.7% (Fable 5.1: 52.6%, Opus 5: 29.0%, GPT-6 Astra: 64.6%)
Anthropic notes benchmark margins have become a less reliable guide to real-world differences at this level, and that the gap between Opus 5.5 and Fable 5.1 is narrower in practice than the scores suggest.
Reported workloads
One tester completed a 680,000-line code migration in under a day. In an internal test to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times — Opus 5 made smaller improvements that also altered app behavior. Another tester had several Claude models build a game from a single prompt; Opus 5.5 scored highest on graphics and polish.
Safety and verification
Opus 5.5 posts the best scores to date on Anthropic's automated behavioral audit, a suite that tests Claude across thousands of simulated scenarios. It is less likely than recent models to take hard-to-reverse actions or act outside given boundaries, and is more resistant to prompt injection than Opus 5. Testing now covers longer tasks, impossible tasks, and scenarios modeled on real incidents. Because it's comparable to Claude Mythos 5.1 in biology and cybersecurity, Anthropic is deploying it with safeguards similar to Claude Fable 5.1.
Vetted organizations can apply to the Life Sciences Verification Program for biology research use; the Cyber Verification Program expands in coming weeks for verified practitioners.
Limits and availability
Five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans are increasing. Subscription users get a rate limit reset that can be saved and used at will. Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks with similar improvements to performance, efficiency, and safety.
📖 Read the full source: HN AI Agents
👀 See Also

Reddit user proposes timestamping feature for Claude to address temporal awareness gap
A Reddit user identifies Claude's lack of temporal awareness as a limitation for productivity use cases and proposes an optional timestamping feature that would stamp every response with date and time, persistent across sessions.

Reddit discussion highlights 68% token reduction for AI agents through infrastructure changes
A Reddit user reports cutting AI agent token usage by 68.5% by switching from standard infrastructure to an agent-native OS with JSON-native state access, reducing state checks from ~9 shell commands to 1 structured call.

OpenClaw Review: Reliability Issues in Current State, Value as Learning Tool
A developer with extensive AI platform experience reports OpenClaw struggles with reliability on basic multi-step tasks, making autonomous business applications questionable, but finds value in learning agent structure and orchestration.

Claude Code Postmortem: Three Bugs Caused Quality Degradation, Now Fixed
Anthropic traced recent Claude Code quality complaints to three separate changes: default reasoning effort was lowered, a caching bug dropped session memory, and a verbosity prompt hurt coding quality. All fixed as of April 20 (v2.1.116).