Local LLM Benchmark: Backend Generation by Function Calling – GLM, Qwen, DeepSeek Compared

Five months after an initial uncontrolled measurement, AutoBe.dev has published a proper benchmark of local and frontier LLMs for backend code generation using function calling. The benchmark uses a controlled variable setup with a real scoring rubric, testing models on generating recursive-union AST schemas via a function calling harness.
Key Findings
- The function calling harness has effectively closed the gap between frontier and local models on backend generation. Specifically,
gpt-5.4's DB/API design scores are approximately equal toqwen3.5-35b-a3b, andclaude-sonnet-4.6's logic scores matchqwen3.5-27b. - This is the last round including frontier models. Running them monthly costs ~200–300M tokens (~$1,000–$1,500 per model on GPT 5.5 pricing). From next month, only OpenRouter endpoints under $0.25/M tokens or models that fit on a 64GB unified-memory laptop will be included.
- Frontend automation will be added to the benchmark in the June/July round, using the SDK AutoBe already emits to drive end-to-end AI-built frontends (visuals rough, but all functions work).
Unexpected Inversions
Several results are still under investigation:
openai/gpt-5.4scores below its ownminisibling.deepseek-v4-prolands one notch belowqwen3.5-35b-a3band barely separates from its ownFlashsibling.- Within the Qwen family, dense 27B beats every MoE variant, including 397B-A17B.
Possible explanations being investigated include CoT-compliance phenomenon (larger/frontier models tend to skip procedural instructions enforced by the harness) and benchmark defects (n=4 reference projects, narrow score band, harness scoring own pipeline).
Recommended Models
Three locked-in candidates for next month:
openai/gpt-5.4-nano— $0.25/M tokensqwen/qwen3.6-27b— $0.195/M tokensdeepseek/deepseek-v4-flash— $0.14/M tokens
All are under $0.25/M on OpenRouter or runnable on a 64GB unified-memory laptop, and handle function calling cleanly.
References
- Benchmark Dashboard: https://autobe.dev/benchmark/
- Generation Results: GitHub: autobe-examples
- GitHub Repository: https://github.com/wrtnlabs/autobe
📖 Read the full source: r/LocalLLaMA
👀 See Also

Decoupled DiLoCo: Resilient Distributed Training Across Data Centers with Low Bandwidth
Google DeepMind's Decoupled DiLoCo trains LLMs across distant data centers using 2-5 Gbps WAN, with self-healing islands of compute that isolate hardware failures without degrading ML performance.

AI Coding Agents Can Fragment Workflow and Drain Attention, Developer Warns
A 12-year web dev reports that using Claude Code daily leads to micro interruptions, loss of focus, and mental exhaustion — without measurable productivity gains.
Parameter Golf: OpenAI's AI-Assisted ML Research Experiment
OpenAI ran Parameter Golf, a competition with 1,000+ participants and 2,000+ submissions, testing AI-assisted machine learning, coding agents, quantization, and novel model design under strict constraints.

OpenClaw's Context Management Criticized as Token-Intensive and Architecturally Flawed
A Reddit post criticizes OpenClaw for inefficient context handling that leads to excessive token usage. The framework appends all actions to global history, creating bloated prompts that overwhelm smaller models and force reliance on expensive frontier models like Claude Opus.