Open-source local hook automatically switches Claude models to cut AI costs

A developer has open-sourced a local hook that automatically selects the most cost-effective Claude AI model based on the type of coding task, potentially reducing AI costs by 50-70% without quality loss.
How it works
The tool runs as a local hook in Cursor and Claude Code (both use the same hook system) before each prompt is sent. It sits next to Opus/plan and acts as an efficient front-end filter that prevents obviously bad model matches before they hit expensive models.
Key functionality
- Reads the prompt and current model selection
- Uses simple keyword rules to classify tasks (git operations, feature work, architecture/deep analysis)
- Blocks if you're overpaying (e.g., Opus for git commit) and suggests Haiku or Sonnet
- Blocks if you're underpowered (Sonnet/Haiku for architecture) and suggests Opus
- Lets everything else through unchanged
- ! prefix bypasses the filter completely if you disagree with its suggestion
Technical details
- 3 files: bash + python3 + JSON
- No proxy, no API calls, no external services
- Fail-open design: if it hangs, Claude Code proceeds normally
- Open-sourced at: https://github.com/coyvalyss1/model-matchmaker
Performance and testing
The developer analyzed several weeks of their own prompts and found:
- 60-70% were standard feature work Sonnet could handle
- 5-20% were debugging/troubleshooting
- A significant portion were pure git/rename/formatting tasks that Haiku handles identically at 90% less cost
Retroactive analysis showed the tool would have cut 50-70% of AI spend with no quality drop. After tuning, it correctly handled 12/12 real test prompts.
Problem it solves
The issue isn't knowledge—developers know they should switch models—but friction. When in flow state, developers don't want to think about dropdown menus. This tool automates the decision-making process.
📖 Read the full source: r/ClaudeAI
👀 See Also

Prefex: A Local Proxy for Claude Code That Automates Prompt Caching and Session Memory
Prefex is a local proxy that sits between Claude Code and Anthropic's API, automatically injecting the header required for Anthropic's beta prompt caching feature. It also implements session memory to avoid resending full conversation history and includes a model router for cost optimization.

HyperResearch: Open-source Claude Code skill harness turns it into a deep research agent
HyperResearch converts Claude Code into a 16-step deep research pipeline with persistent knowledge store, fact-checking, and authenticated web sessions. Open-source, single-command install, outperforms OpenAI and Google on DeepResearch Bench.

Custom Voice Extraction Process for Claude Code with Template
A developer shares a three-pass extraction process to create a custom voice skill for Claude Code, resulting in a 510-line SKILL.md file with ban lists for LLM-isms, anti-performative rules, and format-specific voice modes. The open-source template works with any language using 10+ writing samples.
ThoughtDAG: An Editable Context Graph for LLM Conversations
ThoughtDAG turns LLM context into an editable graph, letting you branch, merge, and prune what the model sees. A pilot shows it can fix errors by removing outdated branches, reducing token counts.