Omnicoder-9B Performance Review: Speed vs. Tool Calling Issues

Technical Overview
Omnicoder-9B is a coding-specific model developed by Tesslate, based on the Qwen 3.5 architecture. It's fine-tuned on top of Qwen3.5 9B using outputs from multiple models including Opus 4.6, GPT 5.4, GPT 5.3 Codex, and Gemini 3.1 Pro.
Performance Characteristics
The model demonstrates strong performance on mid-tier hardware. With 12GB of VRAM, users report consistent token generation at 15 tokens/second even with context size set to 100k. Prompt processing is notably fast at approximately 265 tokens/second. The model runs without crashing systems or causing performance degradation.
Limitations and Issues
Despite the speed advantages, Omnicoder-9B shows several limitations in practical coding scenarios:
- Failed to generate a complete Super Mario clone in a standalone HTML file with a one-shot prompt
- Experienced tool calling failures with MCP servers, generating MCP errors during data fetching
- Issues executing write tool calls from Claude Code, though this may involve compatibility factors
IDE Integration Testing
Testing in development environments revealed mixed results:
- In LM Studio with Roo Code: Disconnections occurred as token size increased to 4k, though this appears to be an integration issue rather than model-specific
- The model successfully updated or wrote small scripts with token sizes between 2-3k
- API requests failed for tokens above 4k without error messages
- In Claude Code: Token generation felt slower compared to Roo Code, and the model failed to execute write tool calls after generating output
The user notes that Roo Code has been the most effective extension for local LLMs among Continue and other tested options.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Testing MiniMax M2.7 via API on Three Real ML and Coding Workflows
A developer benchmarks MiniMax M2.7 against Claude Opus 4.7 on three real tasks: refactoring a PyTorch project, drafting Obsidian notes, and more. Key findings and setup included.

Claude Code Session Dashboard: Open Source Tool for Monitoring Multiple Sessions
An open-source dashboard that monitors multiple Claude Code sessions simultaneously, showing token usage, costs, session status, context window usage, and active subagents. Installation requires three commands: git clone, cd, and npm install && npm start.

Using Obliteratus toolkit to remove refusal weights from AI models
A Reddit user used the Obliteratus toolkit to surgically remove specific weights responsible for refusal behavior in AI models, demonstrating on Alibaba's Qwen 1.5B model that it can reveal training origins without retraining.

HolyClaude: Docker Container for Claude Code with Browser UI and Headless Chromium
HolyClaude is an open-source Docker container that packages Claude Code CLI with a browser UI, headless Chromium, and additional AI coding tools. Setup requires only docker compose up and provides access at localhost:3001.