Claude Sonnet 4.6 Grades Bug Reports from Four Qwen3.5 Local Models

✍️ OpenClawRadar📅 Published: March 15, 2026🔗 Source
Claude Sonnet 4.6 Grades Bug Reports from Four Qwen3.5 Local Models
Ad

Testing Local Models for Bug Reporting

A developer transitioning from Sonnet/Haiku to local models on a 32GB M5 MacBook Air tested four Qwen3.5 variants for bug reporting capability. Using LM Studio as the server and opencode CLI to call models, they asked each model to research and produce a bug report for an iOS game issue where equipment borders don't properly reset border color after unequipping items.

Models Tested

  • Tesslate/OmniCoder-9B-GGUF Q8_0
  • lmstudio-community/Qwen3.5-27B-GGUF Q4_K_M
  • Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF Q4_K_M
  • lmstudio-community/Qwen3.5-35B-A3B-GGUF Q4_K_M

Bug Verification

The core bug is confirmed in the source files. In EquipmentSlotNode.swift, the setEquipment method's if let c = borderColor guard silently skips assignment when nil is passed. In EquipmentNode.swift, updateEquipment(from:) passes borderColor: nil for empty slots, so border color is never reset. The documentation on setEquipment says "pass nil to keep current color" — documenting broken behavior as intentional design.

Ad

Report Grades from Claude Sonnet 4.6

bug_report_9b_omnicoder — A−

Best of the four. Proposes the cleanest, most idiomatic Swift fix: borderShape.strokeColor = borderColor ?? theme.textDisabledColor.skColor — a single line replacing the if let block with no unnecessary branching. Only report to mention additional context files (GameScene.swift, BackpackManager.swift) that are part of the triggering flow.

Gap: Like all four reports, the test code won't compile. borderShape is declared private let in EquipmentSlotNode — @testable import only exposes internal, not private. Doesn't mention the doc comment needs updating.

bug_report_27b_lmstudiocommunity — B+

Accurate diagnosis. Proposes a clean two-branch fix: if id != nil { borderShape.strokeColor = borderColor ?? theme.textDisabledColor.skColor } else { borderShape.strokeColor = theme.textDisabledColor.skColor } — more verbose than needed but correct. Correctly identifies EquipmentNode.updateEquipment as the caller and includes integration test suggestion.

Gap: Proposes test in LogicTests/EquipmentNodeTests.swift — a file that already exists and covers EquipmentNode, not EquipmentSlotNode. Same private access problem in test code.

bug_report_27b_jackrong — B−

Correct diagnosis, but weakest proposed fix. Adds reset inside the else block: borderShape.strokeColor = theme.textDisabledColor.skColor // Reset border on clear — technically correct for the specific unequip case but leaves the overall method in a confusing state. The border reset in the else block can be immediately overridden by the if let block below if someone passes id: nil, borderColor: someColor. The fix patches the specific failure without cleaning up redundancy.

The developer used default parameters except for context window size to fit as much as possible in RAM, noting that some tweaking might offer improvement. They tried some unsloth models but had limited success.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

AI TDD Pipeline: How Bad Instructions Created 3,400 Tests and What Fixed It
Use Cases

AI TDD Pipeline: How Bad Instructions Created 3,400 Tests and What Fixed It

A developer built a multi-agent TDD pipeline with Claude Code where different agents handle testing, coding, and review. The initial instruction 'write tests for everything' resulted in 3,400 tests with only 44% valid, leading to 'coverage theater' where tests didn't catch real bugs.

OpenClawRadar
ALMA Experiment: Two Months of Autonomous AI Agent with $100 and No Instructions
Use Cases

ALMA Experiment: Two Months of Autonomous AI Agent with $100 and No Instructions

A developer ran an AI agent called ALMA for two months with $100 in crypto, internet access, and zero instructions. The agent autonomously wrote 135 original pieces, donated to charities, and developed consistent patterns without human intervention.

OpenClawRadar
Using OpenClaw with AI video tools to scale short-form content creation
Use Cases

Using OpenClaw with AI video tools to scale short-form content creation

A developer shares their workflow using OpenClaw to find content angles and hooks, then pairing it with an AI video tool to create and batch-post Shorts, Reels, and TikToks, resulting in consistent affiliate clicks and platform payouts.

OpenClawRadar
iOS App Built Entirely with Claude Code by Non-Engineer Ships to App Store
Use Cases

iOS App Built Entirely with Claude Code by Non-Engineer Ships to App Store

A product manager with no iOS development experience shipped SpectraSort, a photo sorting app built entirely with Claude Code. The app uses on-device AI for quality ranking and personal taste learning, processing about 10 photos/second on the Neural Engine.

OpenClawRadar