Mnemos: an MCP server for persistent Claude Code memory

Claude Code forgets everything between sessions, forcing you to re-explain conventions, corrections, and context every time. Mnemos (GitHub) is an MCP server that fixes that by storing and replaying persistent memory across sessions.
How it works
- On session start, it pushes a ranked context block (conventions, corrections, skills, hot files, recent session summaries) into Claude's prompt.
- Records corrections as
tried / wrong_because / fix. Three corrections on the same topic auto-promote into a reusable skill withWhen this applies / Avoid / Dosections — deterministic pattern mining, no LLM in the loop. - Bi-temporal store: facts carry valid/invalid timestamps, so "we used to use X, now Y" works without stale context.
- Compaction recovery: one tool call restores the goal and key decisions after Claude Code compacts mid-session.
- Prompt-injection scanner at write boundary (instruction overrides, zero-width unicode, MCP spoofing).
- Retrospective replay: regenerate any past session as markdown with everything learned since layered in, paste it back to Claude, ask "what would I do differently now."
Stack & install
- Single static Go binary, 15 MB. No Python, no Docker, no vector DB, no CGO.
- SQLite + FTS5 retrieval, optional cosine similarity if Ollama is running.
- Install (MIT, free, no paid tier):
curl -fsSL https://raw.githubusercontent.com/polyxmedia/mnemos/main/scripts/install.sh | bash
mnemos init
mnemos initauto-wires Claude Code, Claude Desktop, Cursor, Windsurf, and Codex CLI. Restart your agent andmnemos_*tools show up.
Who it's for
Developers using Claude Code who are tired of re-teaching conventions every session and want reproducible, token-free memory.
📖 Read the full source: r/ClaudeAI
👀 See Also

Holisto Seed: A Local LLM Framework with Persistent Identity and Consensual Memory Consolidation
Holisto Seed is a Relational Individuation Framework that gives LLM agents persistent identity, biographical memory, and co-evolutionary relationships with users. It runs fully local with a Git-based versioning system and features a consensual sleep cycle for memory consolidation.

Agent frameworks waste 350,000+ tokens per session resending static files
A benchmark on a local Qwen 3.5 122B setup revealed agent frameworks waste over 350,000 tokens per session by resending static files. A compile-time approach reduced query context from 1,373 tokens to 73, achieving a 95% reduction.

OnPrem.LLM AgentExecutor: Launch Sandboxed AI Agents with Built-in Tools
OnPrem.LLM's AgentExecutor lets you create autonomous AI agents that execute complex tasks using cloud or local models, with nine built-in tools including file operations, shell commands, and web search. You can run agents in sandboxed containers for security.

Claude Desktop Feature Request: Session Start Hook for Automatic Initialization
A developer building persistent context systems for Claude Desktop identifies a gap: the User Preferences field only injects instructions when the user sends the first message, requiring manual triggers for initialization. They propose adding an "On Session Start" execution field that runs automatically when a new conversation opens.