PRECC Tool Cuts Claude Code API Costs with Pre-Tool-Call Compression

PRECC is an open source tool that reduces Claude Code API costs by compressing redundant context before it reaches the model. It uses a pre-tool-call hook that intercepts Bash, Read, and Grep calls to apply compression algorithms.
How It Works
The tool addresses cost issues where API bills were climbing due to redundant context being sent multiple times. Common sources of waste include:
- Same file contents sent repeatedly
- Verbose shell output
- Overlapping grep results that the model doesn't need in full
The pre-tool-call hook runs RTK (Redundancy-aware Token Kompression) on tool output before it reaches Claude. The compression process:
- Deduplicates repeated spans
- Strips noise
- Summarises large reads
- Returns compressed version to the model
Performance Results
The hook runs in approximately 2.93ms, adding no perceptible latency to operations. In practice, users see 40-66% fewer input tokens across typical coding sessions. Model output quality remains unchanged because the compression preserves signal while stripping redundancy.
This type of optimization is particularly useful for developers using Claude Code extensively, where repeated file reads and tool outputs can significantly increase token usage and costs.
📖 Read the full source: r/ClaudeAI
👀 See Also

2026 Hermes Agent Alternatives Roundup: Self-Hosted Options from OpenClaw to memU Bot
A developer who has been running Hermes since launch tested every self-hosted and managed alternative after the ClawHub security mess. Key findings: OpenClaw (370k stars) but 9 CVEs in 4 days and ~20% malicious packages; TrustClaw rebuilt with OAuth/sandboxing; nanobot at ~4K lines Python with MCP; memU Bot with unique structured memory. Managed options include Perplexity Computer (19 models, $200/mo), Claude Cowork (opens real Mac apps), and KimiClaw (40GB RAG, locked to K2.5, Chinese data law). Full roundup at source.

Microsoft BitNet: 1-bit LLM inference framework for CPU and GPU
Microsoft released BitNet, an inference framework for 1-bit LLMs that achieves 1.37x to 6.17x speedups on CPUs and reduces energy consumption by 55.4% to 82.2%. It can run a 100B parameter model on a single CPU at 5-7 tokens per second.

RunAnywhere RCLI: On-Device Voice AI Pipeline for Apple Silicon
RunAnywhere has released RCLI, an open-source voice AI pipeline for macOS that runs STT, LLM, and TTS entirely on Apple Silicon devices. The tool uses their proprietary MetalRT inference engine and claims significant performance improvements over existing solutions.

Open-source Claude Code skill /unzuck curates social media feeds into dashboard
A free, open-source Claude Code skill called /unzuck scans feeds across Hacker News, Reddit, LinkedIn, YouTube, Twitter/X, Instagram, and Facebook in parallel using browser automation, scores items against user interest profiles, and generates interactive HTML dashboards.