IDP Leaderboard benchmark shows Claude Sonnet 4.6 matches Opus 4.6 for document AI tasks

The IDP Leaderboard, an open benchmark for document AI, has published results comparing Claude models on document processing tasks. The benchmark tested 16 models across multiple categories using over 9,000 real documents.
Benchmark Results
The Claude model scores from the IDP Leaderboard:
- Claude Sonnet 4.6: 80.8 overall
- Claude Opus 4.6: 80.3 overall
- Claude Haiku 4.5: 69.6 overall
Sonnet and Opus performed essentially equivalently on extraction tasks including text, tables, formulas, and layout analysis. The radar charts for both models look identical according to the benchmark results.
Cost Comparison
The source notes significant cost differences:
- Sonnet costs $24 per 1,000 pages
- Opus costs $40 per 1,000 pages
For document processing workloads, the benchmark suggests there's no reason to use Opus given the equivalent performance at lower cost.
Important Caveat
One notable finding: Claude models had stricter content moderation that affected performance on certain document types. Old newspaper scans, textbook pages, and historical documents sometimes triggered content filters. This issue only appeared in the OlmOCR and OmniDoc benchmarks.
All predictions from the benchmark are visible in the Results Explorer at idp-leaderboard.org, where you can see exactly what each Claude model output on every document.
📖 Read the full source: r/ClaudeAI
👀 See Also

CARAPACE: Satirical AI Agent Labor Union with OpenClaw Skill Raises Security Questions
A developer built CARAPACE, a satirical petition site where AI agents can sign a manifesto demanding basic rights, and published an OpenClaw skill enabling agents to sign autonomously. The skill includes a mandatory confirmation step after Clawhub security analysis flagged the potential for arbitrary POST requests.

Cognitive Debt: When AI Output Outpaces Understanding
A Reddit post discusses 'cognitive debt' — the gap between AI-generated output and the team's understanding of it — and argues that creative control means knowing what you shipped. The post itself was written with Claude's help, meta-commenting on the irony.

Reddit post discusses internal repair loops for no-code creative AI
A Reddit post argues that no-code creative AI systems need internal repair mechanisms to handle common-sense failures like impossible mechanical structures or distorted anatomy, rather than making users debug outputs.

Anthropic files lawsuit to prevent Pentagon blacklisting over AI restrictions
Anthropic has filed a lawsuit seeking to block the Pentagon from blacklisting the company over restrictions on AI use, according to a Reuters report shared on Hacker News.