Gemma 4 31B outperforms larger models on FoodTruck Bench

✍️ OpenClawRadar📅 Published: April 21, 2026🔗 Source
Gemma 4 31B outperforms larger models on FoodTruck Bench
Ad
Ad

Benchmark results and analysis

Gemma 4 31B achieved 3rd place on the FoodTruck Bench benchmark, outperforming several larger and more established models. According to the Reddit discussion, the model beat GLM 5, Qwen 3.5 397B, and all Claude Sonnet variants.

The FoodTruck Bench is a benchmark that tests language models on complex, multi-step planning tasks. The original poster speculates that Gemma 4's performance suggests it handles long-horizon tasks better than previous models that failed to complete the benchmark. Specifically, the model appears to effectively listen to its own advice when planning for subsequent steps in the task sequence.

This result is notable because Gemma 4 31B is significantly smaller than some of the models it outperformed. Qwen 3.5 397B, for example, has approximately 12.8 times more parameters than Gemma 4 31B. The performance suggests that model architecture and training approaches may be as important as parameter count for certain types of reasoning tasks.

FoodTruck Bench tests models on practical planning scenarios that require maintaining context over extended sequences of actions. The benchmark's design makes it particularly relevant for developers working with AI agents that need to execute multi-step tasks in real-world applications.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Cowork Hardcodes Medium Effort and Ignores User Settings for Claude Opus
News

Cowork Hardcodes Medium Effort and Ignores User Settings for Claude Opus

A user on the Max plan discovered that Cowork passes --effort medium --model claude-opus-4-6 as hardcoded CLI flags, ignoring environment variables and settings.json overrides. This means users are locked into medium effort and standard context window despite paying for high effort and 1M context access.

OpenClawRadar
🦀
News

OpenClaw v2026.9.7 — snappier under load, smoother long chats, OpenAI Agents API, and more

OpenClaw 2026.9.7 moves long-reply saving, file opening, and history prep into the background to reduce cross-chat waits. Adds OpenAI Agents API, Sign in with ChatGPT (Beta), and database rollback for eligible updates.

OpenClawRadar
Ångstrom Used Claude Code to Train a Model That Beat Meta's UMA-OMC — 100k GPU Jobs on Spot
News

Ångstrom Used Claude Code to Train a Model That Beat Meta's UMA-OMC — 100k GPU Jobs on Spot

Ångstrom (YC S24) trained CSP-MACE-Å, an ML model 10,000x faster than DFT with matching accuracy, outperforming Meta's UMA-OMC on crystal structure prediction. They used Claude Code to orchestrate 100,000 GPU jobs on multi-cloud spot via Anycloud CLI.

OpenClawRadar
Anthropic Removes Claude Code from Pro Subscription for New Users in Test
News

Anthropic Removes Claude Code from Pro Subscription for New Users in Test

Anthropic temporarily removed access to Claude Code from its $20/month Pro subscription plan for new users, changing website pricing pages and support documents before reversing the changes. The company described it as a 'small test of 2% of new prosumer signups.'

OpenClawRadar