VibeThinker-3B: A 3B Parameter Model That Matches 671B DeepSeek on AIME Math Benchmarks

A team of nine researchers at Sina Weibo published a 14-page arXiv report this weekend claiming a 3B parameter model—VibeThinker-3B—matches or exceeds reasoning performance of models hundreds of times larger. The model scored 94.3 on AIME 2026 (American Invitational Mathematics Examination), placing it alongside DeepSeek V3.2 (671B parameters) and ahead of Gemini 3 Pro (91.7). With a test-time scaling technique called Claim-Level Reliability Assessment, the score reaches 97.1.
Key Benchmarks
- AIME 2025: 91.4
- AIME 2026: 94.3 (97.1 with CLRA)
- HMMT 2025: 89.3
- BruMO 2025: 93.8
- IMO-AnswerBench: 76.4
- LiveCodeBench v6 (Pass@1): 80.2
- Unseen LeetCode contests (April–May 2026): 96.1% acceptance rate
- IFEval: 93.4
Notably, VibeThinker-3B underperforms on knowledge benchmarks: 70.2 on GPQA-Diamond vs. 91.9 (Gemini 3 Pro) and 87.0 (Claude Opus 4.5). The authors explicitly acknowledge this as consistent with their claim—verifiable reasoning is "parameter-dense," while open-domain knowledge is "parameter-expansive."
The Training Pipeline
VibeThinker-3B is post-trained on Qwen2.5-Coder-3B (Alibaba's Qwen team) using the "Spectrum-to-Signal Principle," a multi-stage pipeline introduced in the team's earlier VibeThinker work. The paper describes a Parametric Compression-Coverage Hypothesis: verifiable reasoning can be compressed into a compact core, while broad knowledge requires more parameters.
Within hours of publication, the paper received 62 upvotes on Hugging Face Daily Papers, the model repo had 130 likes, and the GitHub repo had 685 stars. Skepticism on social media was high—user @orcus108's post accumulating over 161K views asked: "I genuinely don't know if this is a breakthrough or if the benchmarks are broken."
For context: DeepSeek V3.2 has 671B parameters (~224x larger), GLM-5 has 744B, and Kimi K2.5 exceeds 1 trillion. VibeThinker-3B can run on a consumer laptop.
📖 Read the full source: HN AI Agents
👀 See Also
Canvas LMS Assistant via OpenClaw and Composio
A teacher set up a bot that accesses their three Canvas courses through Composio CLI, letting it create Excel student lists, check for broken links, and track page views.
Claude Code v2.1.273 Adds Gateway Hint Headers, MCP Reconnect Alerts, and Remote Control Session Forking
Claude Code v2.1.273 adds opt-in LLM gateway request headers, an MCP disconnect-and-give-up notification, and forking for Remote Control sessions. It also fixes auto-compact firing at roughly half the real context window.

Apple Core AI Framework: First Look at Apple's Emerging AI Agent Foundation
Apple's new Core AI framework documentation page is live, though the content is behind a JavaScript wall. We break down what this means for AI agent development on Apple platforms.
Claude Code v2.1.285 Adds Desktop Handoff, Plugin Config CLI, and Provider Restrictions
Claude Code v2.1.285 ships CLAUDE_CODE_DISABLE_WEB_FETCH, claude --desktop to hand off a session to the desktop app, claude plugin configure for reading and writing plugin options, an allowedProviders managed setting, and a long list of subagent and MCP fixes.