Qwen 3.5 122B MoE at 35 t/s on a Single 3090 with ik_llama.cpp MTP

A developer running a fully local inference stack on a single desktop reports hitting 35 tokens/s on Qwen 3.5 122B MoE using only one 3090, with the key enabler being a fork of llama.cpp that fixes MTP (Multi-Token Prediction) for offloaded experts.
Hardware Config
- AMD 9900X CPU
- 192GB DDR5-5200 RAM (called “the secret weapon”)
- Two 3090s (Ti + standard), no NVLink
Card 1 runs the worker: Qwen3.5-122B-A10B using Unsloth IQ3_S MTP GGUF with 204K context. 75% of expert layers are offloaded to CPU via surgical -ot flags. Card 2 runs the reasoner: Qwen3.6-35B-A3B Q4_K_XL with MTP at 135 t/s, 262K context.
Additional CPU-only instances handle background processing: Dialectic (35B heretic Q8), Scribe-Logos (Gemma4 19B), Moonshot (Gemma4 2B) — totalling ~19GB RAM.
The ik_llama.cpp Finding
Stock llama.cpp’s MTP evaluates each speculated token’s experts sequentially through DDR5, which on reasoning content actually regresses performance — the draft overhead outweighs the acceptance speedup. The ik fork implements fused MoE ops that batch expert reads for speculated tokens, turning MTP from a +4% gain into a +20% gain. The developer reports 35 t/s decode on a 122B model from a single 3090 using this fork.
If you’re offloading experts to RAM on any MoE model, try ik_llama.cpp before giving up on MTP.
Total Build Cost
- ~$1600 for RAM
- ~$1600 for two 3090s
- ~$400 for everything else
- Running cost: electricity only
📖 Read the full source: r/openclaw
👀 See Also

OpenClaw Sub-agents: Don't Treat a Reply as a Completion Receipt
OpenClaw's sessions_spawn is non-blocking — it returns a runId when work is accepted, not complete. A parent can prematurely report success while a child is still running, failed, or lost.

OpenClaw Mega Cheatsheet: Your Gateway to AI Coding Mastery
Dive into the OpenClaw Mega Cheatsheet from r/openclaw—a comprehensive guide packed with essential tips for AI coding and automation enthusiasts.

Three-layer memory architecture for persistent OpenClaw agent context
A developer built a 3-layer memory system on top of OpenClaw's infrastructure to prevent agents from starting each session without context. The architecture includes L1 workspace files injected every turn, L2 semantic memory search, and L3 reference documents opened on demand.

Optimizing GLM-4.7-Flash on M4 Mac Mini with 24GB RAM
A developer shares specific configuration details for running GLM-4.7-Flash on an M4 Mac Mini with 24GB RAM, including Q3_K_XL quantization, 32k context size with MLA, and memory allocation realities for Metal.