Benching local Qwen 3.6 27B as a Codex validator co-agent

A developer on r/LocalLLaMA has been running a local Qwen model beside OpenAI's Codex as a validator and challenger, and built a small reproducible eval suite to quantify which GGUF quant profiles work best for this role. The workflow: Codex handles main repo work; local Qwen challenges the plan, checks for overbuilding, missed hard directives, UI/design issues, bad assumptions, and long-context misses. The author reviews each interaction before proceeding.
Eval suite setup
The suite tests Qwen 3.6 27B GGUF profiles through llama.cpp, including Bartowski and Unsloth variants at different context sizes and KV cache formats (q8, f16). The focus is on real-world failures: missed directives, bad challenge behavior, overbuilding, UI judgment, and long-context misses.
Key findings
- The top-performing profiles on this suite were:
bartowski-128k-f16,bartowski-128k-q8, andunsloth-128k-q8. All three tied on accuracy. - q8 KV cache showed no measured accuracy loss in this specific suite.
- Context size mattered more than f16-vs-q8 KV for this workflow. 65k profiles failed when the suite required >65k tokens.
unsloth-128k-f16loaded but hit memory/throughput pressure on long-context cases on an RTX 5090.
Practical observations
The author reports Qwen is extremely good at catching silent bypasses, overbuilding, and coding-to-completion shortcuts in Codex. For UI-related tasks, Qwen takes the lead in design while Codex implements. The roles reverse: Qwen challenges the plan, and the human reviews before each stage.
Resources
- Project page: https://robert896r1.github.io/qwen-realworld-accuracy-evals/
- Repo: https://github.com/robert896r1/qwen-realworld-accuracy-evals
📖 Read the full source: r/LocalLLaMA
👀 See Also

Zoku: A Tool That Automatically Detects Repeated Workflows in Claude Code
Zoku is a local tool that hooks into Claude Code's event system to record tool actions across sessions, identifies repeated workflow patterns, and then informs Claude about these patterns so it can proactively suggest or execute them. It requires no configuration, has no dependencies, and stores everything locally in ~/.zoku/.

Shipshots MCP Server: Claude Designs App Store Screenshots and Preview Videos
Shipshots is a visual editor with an MCP server that lets Claude design marketing materials through tool calls. It generates app store screenshots, animated preview videos, and social media visuals based on text descriptions.

Introducing OneTool MCP: An Open Source Multi-Tool for Developers
OneTool MCP, built using Claude AI, offers developers over 100 tools for tasks like web searches, library updates, and file management without tool tax or context rot.
Claude Code v2.1.239: Cost Estimates Include 1.1x Premium, /claude-api Upgrade, and Major Fixes
Claude Code v2.1.239 adds US-only inference premium to cost estimates, introduces /claude-api upgrade for Python 1.x, and fixes Bedrock proxy, WebFetch memory, and more.