OmniRecall Beta: FAISS-Powered Memory Injection for Cloud LLM Chats

What OmniRecall Does
OmniRecall is a local mitmproxy bypass that intercepts traffic to cloud chat interfaces (tested on DeepSeek). It hacks into the proprietary SSE fragment stream and forces a long-term memory layer onto a system that was designed to be stateless.
Technical Mechanism
- Deep-Packet Parsing: Reconstructs the full assistant reply by tracking real-time patches
- Command Control: Detects [ADD], [UPDATE], [REMOVE], [CLEAR] from the AI's output
- Local Brain: Maintains memory.txt + FAISS index (sentence-transformers MiniLM-L6)
- Context Injection: Top recalled facts get force-fed into your next message as [RECALL: ...]
Current Status & Limitations
This is a beta/experimental release. The developer notes: "This is the closest I've gotten to the dream after weeks of debugging hell. It is buggy. It is experimental. [ADD] is mostly stable, but [SEARCH] is temperamental—if you want perfection, fix it yourself. I've hit my energy limit on this build."
Upstream UI changes will break it. The developer states: "If it breaks, that's on you now."
Requirements & Setup
Potato-PC Requirements:
- CPU only (faiss-cpu + all-MiniLM-L6-v2)
- No local LLM needed — augments the cloud models you already use
- Zero cost, zero API keys, 100% local data isolation
How to Deploy:
pip install mitmproxy faiss-cpu sentence-transformers numpyTrust the mitmproxy CA cert on your OS/browser (run mitmproxy once to generate it). Set system proxy to 127.0.0.1:8080. Then run:
mitmdump -s omnirecall.pyGo to chat.deepseek.com and start feeding it memories.
License Terms
The project uses an aggressively restrictive source-available license:
- No commercial use
- No private forks
- Mandatory public ALTERATIONS.md for any logic changes
- If you port to Claude/GPT-4o/whatever, keep it public per the license
The developer explains: "I've watched too many solo-dev projects get strip-mined, privatized, or turned into paid SaaS while the creator gets zero. This license isn't friendly—it's built to protect the work from exactly those people. If the terms scare you off, that's the point."
📖 Read the full source: r/LocalLLaMA
👀 See Also

Building a Programming Language with Claude Code: The Cutlet Experiment
Ankur Sethi built a complete programming language called Cutlet using Claude Code over four weeks, with the AI generating every line of code while he focused on guardrails and testing. The language features dynamic typing, vectorized operations, and a REPL, running on macOS and Linux.

Hollow AgentOS reduces Claude Code token usage by 68.5% with JSON-native OS for AI agents
Hollow AgentOS is a JSON-native operating system for AI agents that cuts Claude Code's token usage by 68.5% by eliminating wasteful shell command overhead. It plugs into Claude Code via MCP, runs local inference through Ollama, and is MIT licensed.

Claude Code user creates /discuss command for read-only conversations
A Claude Code user created a 25-line custom skill called /discuss that enables read-only conversations without file modifications. The command allows code exploration, research, and discussion while preventing edits, using the --dangerously-skip-permissions flag with built-in safety.

Multi-Agent Content Pipeline for Claude Code with Quality Gates
A developer built a six-agent content pipeline for Claude Code that separates research, writing, editing, and SEO tasks with quality gates between stages. The system halts for manual approval before publishing and allows individual agent re-runs.