llama.cpp Massive Prompt Reprocessing with Coding Agents: Debugging KV Cache and Context Swapping

A developer on r/LocalLLaMA is hitting a serious performance issue with llama.cpp when running long-context coding agents (opencode + pi.dev) via llama-swap. Even with highly similar prompts (LCP similarity often >0.99), the system periodically discards the KV cache and reprocesses 40k+ tokens, causing TTFT of multiple minutes.
Observed Behavior
- Context grows to 50k+ tokens.
- After several normal reuses (e.g.,
prompt eval time = 473 ms / 19 tokens),n_pastsuddenly drops to ~4-5k. - llama.cpp then reprocesses the full prompt:
n_tokens = 4750 prompt eval time = 222411 ms / 44016 tokens. - Cache usage hits 4676 MiB, exceeding the configured limit (2500 MiB).
Current Configuration
llama-server --ctx-size 150000 --parallel 1 --ctx-checkpoints 32 --cache-ram 2500 --cache-reuse 256 -no-kvu --no-context-shiftSuspected Causes
- Cache invalidation due to overflow of
--cache-ramlimit – the log shows 4676 MiB used vs 2500 MiB limit. - Bad KV reuse mechanism when early prompt tokens change (possibly frequent alterations by opencode).
- Insufficient
--ctx-checkpointsor--cache-reusefor the 150k context size.
Recommendations from the Community
The thread is thin on answers so far, but obvious first steps include increasing --cache-ram to match typical usage (e.g., 5000+ MiB), or reducing --ctx-size to stay under the cache limit. Also check if opencode is intentionally mutating prompt prefixes; if so, locking the system prompt or using a fixed prefix could improve reuse.
For developers running similar setups, share your working configs in the source thread.
📖 Read the full source: r/LocalLLaMA
👀 See Also

The Prompt Structure That Fixed Claude AI Summaries of Large PDF Reports
A developer shares how switching from 'summarize this' to role + decision + specific extraction prompts turned Claude's generic summary output into actionable risk flags and concrete action items.

Agent Skills: Stop Writing SOPs, Start Building Boundary Systems
A Reddit post argues that adding more skills or tools to an AI agent makes it more fragile. The solution: minimum complete toolset, maximum boundary clarity.

Check Claude's project summaries into your repo — they're better than human docs
A developer suggests committing Claude-generated project summaries to your repo. They're good enough, take seconds to generate, and can help future readers.

After 3 months of A/B testing 160 Claude prompt codes: the boring takeaways
Samarth built a controlled test rig, ran 160 prompt codes through it, and found that most are placebo, 7 consistently shift reasoning, and stacking 3+ codes confuses the model. Skills files outperform prompt codes for Claude Code.