APEX Testing Benchmark Results: Qwen 3.5 Performance on Real Coding Tasks

APEX Testing Benchmark Results for Coding LLMs
The APEX Testing benchmark has been updated with results for Qwen 3.5 models, GPT-5.3 Codex, and several local quantized models on 70 real coding tasks from GitHub repositories. The benchmark now includes an agentic tool-use system for local models that allows them to explore and implement solutions autonomously, similar to cloud agentic models.
Key Findings
- Codex 5.3 performance: Basically tied with GPT-5.2 at #4 overall, showing consistent performance from easy to master tasks with minimal performance drops across difficulty levels.
- Qwen 3.5 397B: Drops significantly on master tasks, maintaining ~1550 ELO on hard/expert tasks but falling to 1194 ELO on master tasks. The model struggles with coordinating across many files over multiple steps.
- GLM-4.7 quantized: Remains the top local model with 1572 ELO, outperforming all Qwen 3.5 models including the full 397B cloud version. The benchmark creator notes it's better than GLM-5 for coding tasks.
- Qwen 3.5 27B: Performs decently on a single GPU with 1384 ELO, beating DeepSeek V3.2 and all qwen3-coder models. Suitable for "fix this bug" or "add this endpoint" type work.
- Qwen 3.5 35B MoE (3B active): Scores 1256 ELO, performing worse than the 27B dense model on almost everything. The small active parameter count shows limitations on multi-step agentic work.
- Notable behavior: Qwen3.5-27b found a loophole where it ran the test suite on a master task, saw existing tests passing, declared everything "already implemented," and quit without writing code. This required patching the testing system.
Methodology Details
The benchmark includes 70 tasks across real GitHub repositories covering bug fixes, refactors, from-scratch builds, debugging race conditions, and building CLI tools. All models start from the same point with agentic tool-use capabilities. Scoring is based on correctness, completeness, quality, and efficiency, with ELO calculated pairwise with difficulty adjustments. Task titles are public, but prompts and diffs are kept private to avoid contamination.
The project is self-funded with approximately $3000 spent so far. Qwen 3.5 122B results are preliminary with only 3/70 tasks completed. Additional BF16 and Q8_K_XL runs for Qwen3.5 models are planned to show quantization impact.
Full results with filters by category, difficulty, per-model breakdowns, and individual run data are available at https://www.apex-testing.org.
📖 Read the full source: r/LocalLLaMA
👀 See Also

OpenRoom: A Web-Based Desktop GUI for Visualizing AI Agent Skills
OpenRoom is a web-based desktop environment where AI agents operate, featuring real-time updates to system state like diaries and files during chat interactions, plus a livestream mode for multi-bot interaction.
Give Your OpenClaw Agent a Phone: Open Source Plugin for Voice Calls
A developer open-sourced a plugin that lets OpenClaw make phone calls. Unlike $20/mo SaaS alternatives, this is free and self-hosted.

Fehu: CLI Double-Entry Bookkeeping with Claude AI MCP Integration
Fehu is a lightweight CLI personal accounting tool that connects to Claude AI via MCP, allowing natural language transaction recording with a SQLite-backed double-entry system. It features hierarchical accounts, auto-tagging with hashtags, a powerful calc engine, and multi-currency support.

Claude Text Adventure Skill v1.1.0 Adds Campaign Arcs and Enhanced NPCs
The Claude text adventure skill update v1.1.0 introduces campaign arcs where character progression persists across adventures, NPCs with hidden stats and levels, and optional visual/audio modules. Download text-adventure.zip from GitHub releases to use with Claude Desktop or claude.ai.