Route Claude Code through Ollama and Cut Your Bill ~90%

This repo by Coherence Daddy provides a complete setup to route Claude Code terminal sessions through a local Ollama instance while keeping Claude Desktop on Anthropic's paid Pro tier. The result: a claimed ~90% reduction in Claude Code API costs.
How It Works
You run two engines side by side:
- Claude Desktop (Anthropic) – used for strategy, architecture, code review, and tricky bugs.
- Claude Code → Ollama – used for lints, refactors, repetitive edits, batch file ops, and grep-and-replace tasks. Runs on a free open-source model (Gemma, Qwen, DeepSeek, your choice).
Setup Process
The repo includes a self-contained HTML presentation (21 slides) with a copy-paste prompt that does ~98% of the setup automatically. It auto-detects your OS (macOS, Windows + WSL2, Linux), installs everything, configures the router, and verifies both engines at the end.
To run locally:
git clone https://github.com/Coherence-Daddy/use-ollama-to-enhance-claude.git
cd use-ollama-to-enhance-claude/presentation
open index.html # macOS, or drag into browserOr directly use the copy-paste prompt from prompts/copy-paste-prompt.md.
Repository Structure
prompts/copy-paste-prompt.md– the setup prompt.presentation/index.html– full visual deck (no build step required).- Also hosted at coherencedaddy.com/tutorials/use-ollama-to-enhance-claude.
Why This Exists
Claude Pro on desktop is great for thinking and architecture, but Claude Code in the terminal burns through quota fast on context-heavy tasks. Routing those tasks through Ollama (local or cloud-hosted free models) keeps the same UX but at a fraction of the cost.
License
MIT – free to use, fork, or remix.
📖 Read the full source: HN AI Agents
👀 See Also

How an Idle Agent Burned 50M Tokens a Day – and How to Fix It
An idle OpenClaw agent burned 50M tokens a day via heartbeat pings with a bloated session. A Reddit user shares how they traced the leak and fixed it with config changes.

Troubleshooting OpenClaw: A Minimalist Reset Method
A Reddit user shares a five-step method to fix unstable OpenClaw setups by removing all skills, switching to Claude Sonnet, clearing sessions, simplifying SOUL.md, and testing with basic commands.

5 Coherence Checks Before Any OpenClaw Profile Goes Live
Stop chasing perfect spoofing. Internal signal coherence matters more. TLS, locale, WebGL, canvas, and behavior checks from r/openclaw.

Using AI as a Cognitive Partner Instead of a Code Factory
A Reddit post proposes a system prompt called 'Cognitive Authorship Copilot' that forces AI to act as a pair programming partner rather than an autonomous solution generator, with three intervention levels based on task complexity.