Bonsai 2 (Qwen 3.8 27B) as an OpenClaw Fallback: 8GB, 85 tok/s on a 5070
PrismML's Ternary-Bonsai-2-27B is a quantized Qwen3 3.8 27B build that a r/openclaw user is running locally as an OpenClaw fallback model when their ChatGPT and Grok quotas run dry. Their claim: ~8GB footprint, 98% of the base Qwen3 27B's quality retained, and text plus vision support.
Numbers from the setup
- Output speed: 85 tok/s
- Hardware: RTX 5070 with 16GB VRAM, 64GB system memory
- Disk footprint: ~8GB
- Quality retention: reported 98% of Qwen3 3.8 27B
The two GGUFs you need
Files come from the prism-ml/Ternary-Bonsai-2-27B-gguf repo on Hugging Face:
Ternary-Bonsai-2-27B-PQ2_0.gguf (text) Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf (vision)
Both are required if you want the multimodal path — the mmproj file is the vision projector, the PQ2_0 file is the ternarized text weights.
Setup notes
Bonsai 2 requires PrismML's own fork of llama.cpp — it will not load in stock llama.cpp. The author had GPT wire up the fork on a separate AI machine from the OpenClaw host, with the model configured to:
- Dynamically load when OpenClaw calls it
- Unload after 10 minutes of inactivity
- Act as the fallback when GPT 5.6 and Grok hit usage limits
What it actually did
The interesting part isn't the tok/s — it's the agentic behavior. Per the post, it handled web search via a browser, VM control, installing and using new software, and running existing workflows, to the point the author says they "can't tell it's not a frontier model."
That's a single user's report, not a benchmark, so treat the quality claims as anecdotal. The load/unload-on-idle pattern is the practical takeaway: it keeps a local 27B resident only when your paid provider is rate-limited, which is a sensible way to avoid permanently burning VRAM on a box that's also doing other work.
If you have a GPU with 16GB+ VRAM, the PQ2_0 + mmproj pair is small enough to be worth testing against your own agentic tasks before deciding whether it holds up for you.
📖 Read the full source: r/openclaw
👀 See Also

Mike: Open-Source Legal AI with Self-Hosting, Multi-Model Support
Mike is an open-source alternative to Harvey and Legora, offering document chat, tabular extraction, and workflow templates — all self-hostable with your own Claude or Gemini API keys.

ToolLoop: Open-Source Framework for Claude-Style Tools with Any LLM
ToolLoop is an open-source Python framework with 11 tools for file operations, code search, shell access, and sub-agents that works with any LLM through LiteLLM. The 2,700-line framework allows switching models mid-conversation while maintaining shared context.

TRELLIS.2 Image-to-3D Ported to Run Natively on Apple Silicon
A developer has ported Microsoft's 4B parameter TRELLIS.2 image-to-3D model to run natively on Apple Silicon via PyTorch MPS, replacing CUDA-specific operations with pure-PyTorch alternatives. The port generates ~400K vertex meshes from single photos in about 3.5 minutes on M4 Pro with 24GB memory.

Giving Claude a Local LLM as an Assistant via MCP on Mac
A developer connects Claude to a local Qwen 2.5 Coder 14B via Ollama and MCP, creating a no-cost assistant for delegating tasks like text processing and handling large files.