Local vs Cloud Models: Qwen-3.6-27B, Gemma-4-31B, Claude Haiku, Codex-Spark on Hard Code Gen

A Reddit user compared locally-ran Qwen-3.6-27B (GGUF q4_k_m) against API equivalents: Qwen-3.6-27B via OpenRouter, Gemma-4-31B via OpenRouter, Claude Haiku 4.5, and GPT-Codex-Spark. The test involved implementing an autoresearch loop from a design document — a deliberately hard task to evaluate failure cleanliness, not success rate.
Hardware Setup
- CPU: Ryzen 7 7800X3D
- RAM: 64 GB DDR5-6400
- GPU: RTX 5080 (16 GB VRAM)
- Local model: Qwen-3.6-27B q4_k_m (GGUF) — fits 16 GB VRAM via quantization
Results
- Gemma-4-31B (API): Failed completely. Wrote skeleton with mocked modules, no tests, no config files (
__init__.py,requirements.txt,pyproject.toml). Cost: $0.112, 803k context tokens consumed, 21k generated. - Codex-Spark (API): Produced beautiful folder structure and code, but imports were hallucinated. No unit tests. Used 1% of $100/mo Spark limits.
- Claude Haiku 4.5 (API): Detailed implementation but failed on correctness. (Further details truncated in source.)
- Qwen-3.6-27B (local q4_k_m): Not explicitly scored, but user notes quantized inference degrades quality vs full-precision API version.
Context
The user argues that typical local-model evals use trivial tasks (e.g., Snake in HTML) where both local and frontier models succeed, making local models look better than they are. This test used a real work project with a design document; only Codex-Spark produced fully written (but broken) code. The point: local models are not yet ready for complex code generation without substantial fixes.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Claude Code v2.1.147: Pinned Sessions, /code-review, and Dozens of Fixes
Claude Code v2.1.147 introduces pinned background sessions, renames /simplify to /code-review with effort levels and --comment, plus fixes for PowerShell, MCP, Windows, and more.

Open-source LLMs outperform Claude Opus 4.6 in trading strategy generation at lower cost
A Reddit user tested 10 LLMs on generating trading strategies, finding open-source models outperformed Claude Opus 4.6 despite being 10x cheaper. Minimax 2.5 and Gemini 3.1 topped the leaderboard.

Cron auto-update broke OpenClaw due to config validation error
A cron job set up to auto-update OpenClaw encountered a config validation issue with the cliBackends field, causing connection loss. The fix involved removing the problematic section and restarting the gateway.

Kimi K2.7-Code: Open-Source Coding Model with Better Token Efficiency
Moonshot AI released Kimi K2.7-Code, an open-source image-text-to-text model with enhanced token efficiency for coding tasks. Available on Hugging Face with 334 likes and Novita inference support.