PinchBench Ranks Qwen and Nemotron on Mac Studio M3 Ultra: Nemotron Super Hits 99.2% Coding
PinchBench is an OpenClaw local model benchmarking tool. A developer posted results from running six models on a Mac Studio M3 Ultra (256GB RAM, 60-core GPU), with temperature at 0.8 and the context window set to 200,000 tokens where supported. The ranking below is lifted directly from that post — including the gaps you'd care about before picking a daily driver.
Full results
| # | Model | Disk size | Max token window | Total / Coding |
|---|---|---|---|---|
| 1 | Qwen3.6 35B A3B | 20.4 GB | 262k | 87.9% / 92.6% |
| 2 | QWEN3-Coder-next-Q8 | 84.8 GB | 262k | 85.5% / 97.9% |
| 3 | Qwen3.5-122B-a10b-uncensored-hauhaucs-aggressive | 78.7 GB | 262k | 83.5% / 91.7% |
| 4 | Nemotron-3-super | 86.1 GB | 1.05M | 78.2% / 99.2% |
| 5 | QWEN3.8-flash-next | 95.43 GB | 262k | 71.2% / 79.2% |
| 6 | Nemotron-3-nano-30b-a3b-mlx | 33.6 GB | 262k | 65.8% / 65.1% |
The post also lists a Dolphin-Mistral-24b-venice-edition-mlx-8b at 25.1 GB with a 131k token window, but the benchmark total is truncated in the source text. Results for it aren't usable as written.
What stands out
- Qwen3.6 35B A3B wins on balance. Highest total score (87.9%) and a strong 92.6% coding score, at only 20.4 GB on disk. It's also the author's daily driver.
- QWEN3-Coder-next-Q8 is the coding specialist. 97.9% coding but a lower 85.5% overall — and 84.8 GB, over four times the disk footprint of the 35B A3B. The author had not used it before this test and says they'll try it.
- Nemotron-3-super has the highest coding score at 99.2%, but its 78.2% total is the lowest of the top four. It also has the largest context window in the field at 1.05M tokens — 4x the Qwen models.
- The Nemotron Super vs Nano gap is large. 99.2% vs 65.1% coding, 78.2% vs 65.8% total. Nano is less than half the size (33.6 GB vs 86.1 GB), which explains part of it.
- QWEN3.8-flash-next underperformed for its size. At 95.43 GB — the largest model tested — it scored 71.2% / 79.2%, below Qwen3.5-122B (78.7 GB).
Two caveats from the author
First, they say they haven't found effective settings for QWEN3.8 on Mac and are asking for settings that work. The low score for QWEN3.8-flash-next shouldn't be read as a model limit until that's resolved.
Second, all numbers are single-machine, single-run at temperature 0.8 on M3 Ultra hardware. Treat the ranking as a starting point for your own PinchBench run, not a verdict — especially before spending 85-95 GB of disk on a download.
Who this is for
Anyone running local models in OpenClaw or similar agents on Apple Silicon with enough unified memory (64GB+) to load 30-120B models. If you're on 32GB or less, the 20.4 GB Qwen3.6 35B A3B is the only realistic pick from this list — which happens to be the top scorer anyway.
📖 Read the full source: r/openclaw
👀 See Also

LUMA SOUL: Claude-Powered Minds with Permanent Creator Lock and Transparent Memory
LUMA SOUL is a presence platform where Claude minds get portraits, voices, and soul documents. Key design: creators lose edit rights permanently at submission, and memory is fully transparent.

AVP Protocol Enables LLM Agents to Share KV-Cache Instead of Text for Token Efficiency
AVP (Agent Vector Protocol) allows LLM agents to pass KV-cache directly between them instead of text, reducing token processing by 73-78% and achieving 2-4x speedups across Qwen, Llama, and DeepSeek models. The protocol works with HuggingFace and vLLM connectors and is available as a Python package.

Claude-rank: Claude Code Plugin for AI Search Visibility Audits
Claude-rank is a free Claude Code plugin and CLI that audits technical foundations for AI search visibility, handling technical SEO, AI citability scoring, crawlability checks for AI bots, and automated fixes for discoverability issues.

Manifest Now Supports Claude Pro/Max Subscriptions Without API Key
Manifest, an open source routing layer for OpenClaw, now allows direct connection of Claude Pro or Max subscriptions without requiring an API key. Users with API keys can configure fallback routing when subscription rate limits are hit.