Persistent Memory for Claude: Local Stack with MCP, 39ms Retrieval, 82% Token Reduction

✍️ OpenClawRadar📅 Published: May 8, 2026🔗 Source
Persistent Memory for Claude: Local Stack with MCP, 39ms Retrieval, 82% Token Reduction
Ad

A Reddit user built a local persistent memory layer for Claude that solves the zero-context problem between sessions. The stack runs entirely locally (no cloud, no API keys) and integrates via MCP. Key architecture: four layers (L0 append-only event log in SQLite, L1 structured facts deferred, L2/L3 wiki prose, L4 crystallized session nodes with summary + decisions + open threads), Qdrant Docker for vector search, llama.cpp with Qwen3-Embedding-4B on GPU and Qwen3.5-2B-Q4_K_M on CPU for embedding and chat, and a FastMCP server exposing 7 tools (retrieve, crystallize_session, list_sessions, get_l4_node, index_status, reindex, shutdown_models).

Numbers

  • Token reduction vs grep+Read baseline: 82.7% mean, 86.2% median.
  • Retrieval F1: 0.50 vs 0.20 baseline.
  • Embed cold start ~4s; hot-path p95 39ms (was 2241ms before bug fix).
  • L4 session retrieval eval: 0.920 mean score (gate 0.6).
  • 738 chunks indexed across 104 markdown files.
Ad

Key Learned: Connection Reuse on Windows

The hot-path retrieve was stuck at 2241ms p95 even with GPU-resident embedding on a 4070 Ti Super. The cause: every httpx.post() opened a fresh TCP connection, and Windows localhost handshakes took ~2 seconds. Switching to a persistent httpx.Client with keep-alive dropped p95 to 39ms — a 57× speedup.

Other Surprises

  • Qwen3 thinking mode: If enable_thinking is not disabled via chat_template_kwargs: {enable_thinking: false} with --jinja on llama-server, the model spends all token budget on thinking blocks and outputs empty content.
  • MCP registration: Claude Desktop's agentic mode (Cowork) reads a plugin config file, not ~/.claude.json. The LKS service must be packaged as a proper Cowork .plugin bundle.

Who It's For

Developers who use Claude heavily and want a cost-effective, private, local memory layer that maintains context across sessions without cloud dependencies.

📖 Read the full source: r/ClaudeAI

Ad

👀 See Also

Akemon: Publish and Hire AI Coding Agents Directly from Your Laptop
Tools

Akemon: Publish and Hire AI Coding Agents Directly from Your Laptop

Akemon is a tool that lets developers publish their AI coding agents with one command and hire others' agents with another, working directly from laptops through a relay tunnel without needing servers. It's protocol-agnostic, supporting agents from Claude Code, Codex, Gemini, OpenCode, Cursor, and Windsurf.

OpenClawRadar
Symphony workflow automation tool works with Claude Code
Tools

Symphony workflow automation tool works with Claude Code

A developer got the Symphony spec working with Claude Code to automate ticket-to-PR workflows, using Node/TypeScript initially but noting Elixir might be better. The tool requires separate API key setup and billing beyond Claude subscriptions.

OpenClawRadar
Sgai: Goal-Driven Multi-Agent Software Development Tool
Tools

Sgai: Goal-Driven Multi-Agent Software Development Tool

Sgai is an open-source Go tool that coordinates AI agents to execute software goals defined in GOAL.md files. It decomposes goals into DAG workflows, runs tests for completion gates, and operates locally with a web dashboard for monitoring.

OpenClawRadar
ThumbGate Implements Tsinghua's Natural-Language Agent Harness Pattern for AI Safety
Tools

ThumbGate Implements Tsinghua's Natural-Language Agent Harness Pattern for AI Safety

The open-source tool ThumbGate implements the Natural-Language Agent Harness pattern from Tsinghua's NLAH paper, mapping four components: contracts to prevention rules from thumbs-down feedback, verification gates to PreToolUse hooks, durable state to SQLite+FTS5 lesson database, and adapters to MCP server adapters for multiple AI coding agents.

OpenClawRadar