Claude Opus 4.1 scores 17.75% on SWE-Bench Pro's private dataset, highlighting memorization vs. reasoning gap

Benchmark results show significant performance gap
Claude Opus 4.1 achieved 80%+ on SWE-Bench Verified, but scored only 17.75% on SWE-Bench Pro's private dataset. This dataset contains 276 tasks from 18 proprietary startup codebases that have never been on GitHub, specifically designed to eliminate data contamination through GPL-licensed public repositories.
Other model results on the same private dataset: GPT-5.2 scored 23.81% (topping the leaderboard) and Gemini 3 Pro scored 17.95%.
Trajectory analysis reveals memorization behavior
Scale AI's analysis found that during testing, models could identify correct file paths to modify before fully reading problem descriptions on familiar repositories. This indicates they were navigating by memory rather than reasoning through the problems.
The 80% score on SWE-Bench Verified was real, but measured a different capability than most people assumed - primarily memory of training data rather than reasoning about novel code.
Practical implications for AI coding tool deployment
For developers deciding where to deploy AI coding tools in their workflow, the distinction between memory and reasoning matters more than headline benchmark numbers. Models that perform well on contaminated benchmarks may struggle with truly novel codebases they haven't seen during training.
SWE-Bench Pro was created specifically to address this contamination issue by using code that has never been publicly available on GitHub or in training datasets.
📖 Read the full source: r/ClaudeAI
👀 See Also

Wikipedia Bans AI-Generated Content, Allows Limited AI Use with Human Review
Wikipedia has officially banned its 260,000 editors from using AI like ChatGPT to write articles, citing accuracy and reliability concerns. Editors can still use AI for translation and copy editing with human approval.

Polsia Platform Shows Repetitive SaaS Patterns in Live Founder Launches
Polsia is an autonomous business platform where users describe their business, pay money, and it executes autonomously. A behavioral scientist observed 72 hours of live founder launches, identifying repetitive patterns like AI SDR automation solutions and underserved international markets.

Four UX/Product Gaps Identified in Claude's Onboarding Experience
A user identified four specific UX/product gaps while setting up Claude across Desktop, Cowork, Dispatch, and the iPhone app during active use. Issues include Dispatch tasks entering infinite loops when desktop is offline, single persistent threads in Dispatch, tab-anchored chat panels in Chrome, and missing Google Drive files in the mobile app knowledge base UI.
Markdown Memory Beats a 382-Dependency Memory Runtime in Agent Recall Test
A developer tested a vendor memory runtime against a folder of Markdown files. The runtime returned a superseded API decision; the Markdown folder returned the current one, with sources, and ran for seven months.