agentcache: Python Library for Multi-Agent LLM Prefix Caching

agentcache is a Python library designed to optimize multi-agent LLM systems by implementing prefix caching as a core feature. The library addresses the common problem where frameworks like CrewAI, AutoGen, and open-multi-agent create fresh sessions for each worker, resulting in zero cache hits and duplicated prompt costs.
How It Works
The library operates on a fork-based approach instead of creating separate sessions:
- Start one session with a shared system prompt
- Make the first call - provider computes and caches the prefix
- When you need N workers, fork instead of creating N new sessions
- Parent session: [system, msg1, msg2, ...]
- Forked session: [system, msg1, msg2, ..., WORKER_TASK]
- Exact same prefix = cache hit
Key Features
- Cache-safe forks: Maintains identical prefixes across worker sessions
- Cache-break detection: Diffs snapshots and reports exactly what changed when cache hits drop
- Cache-safe compaction: For long-running sessions, scans old tool outputs before each call and replaces large results with deterministic placeholders to maintain smaller context while preserving cacheable prefixes
- Parameter freezing: Freezes cache-relevant parameters before forking (system prompt, model, tools, messages, reasoning config)
- Task DAG scheduling: Enables parallel workers from one cached session
Performance Results
In a head-to-head test with GPT-4o-mini (coordinator + 3 workers, same task):
- Text injection / separate sessions: 0% cache hits, 85.7 seconds
- Prefix forks: 75.8% cache hits, 37.4 seconds
- Per worker cache hit rates typically range from 80-99%
Installation and Usage
Install via pip:
pip install "git+https://github.com/masteragentcoder/agentcache.git@main"
The library is available on GitHub at github.com/masteragentcoder/agentcache.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Quell Proxy Fixes Claude Code Scroll-Jumping on Windows
Quell is a Rust proxy that sits between your terminal and Claude Code, stripping clear-screen sequences that cause scroll position resets during long responses. It also adds Shift+Enter for newlines, security filtering, and full Unicode support.

V6rge AI Suite Update Adds NVIDIA GPU Support and Beta Coding Agent
V6rge AI Suite has released an update that fixes GPU detection issues, adds full NVIDIA GPU support for better performance, and introduces a new beta coding agent that generates and assists with code directly inside the app.

JobPilot: Claude Code Plugin for Automated Job Applications
JobPilot is a Claude Code plugin that automates job searching and application processes using Playwright browser automation. It includes commands for searching job boards, auto-filling applications, generating cover letters, and tracking application statistics.

OpenHelm: A Local Background Scheduler for Claude Code with Self-Correcting Retry Logic
OpenHelm is a Tauri-based application that runs Claude Code tasks in the background on a schedule, stores all state locally in SQLite, and includes a self-correcting retry loop that adjusts prompts after failures.