Improving Claude Code Sessions with claude-self-improve

claude-self-improve is a command-line tool created to address repetitive errors in Claude Code sessions by automating the analysis and updating of memory files. This tool was developed due to the tedious nature of manually curating MEMORY.md files to improve AI-driven coding sessions.
The process involves three main steps:
- The tool reads session facets, specifically the JSON performance data Claude Code already generates.
- This data is sent to headless Claude (also referred to as Sonnet), where it extracts patterns indicating friction and successes, as well as lessons learned.
- Upon analysis, it updates
MEMORY.mdautonomously, making sure the subsequent sessions are incrementally smarter.
After evaluating 52 sessions, the tool reported a 42% friction rate and identified common anti-patterns. For instance, 'wrong initial diagnosis' accounted for 41% of the friction events. The system also suggested four memory updates and three CLAUDE.md improvements without the need for manual review.
# Example command
bash claude-self-improve.shRunning the tool costs approximately $0.07 to $0.20 per session. The code repository is available on GitHub, providing access to the script and allowing other developers to implement similar improvements.
📖 Read the full source: r/ClaudeAI
👀 See Also

Qwen3.6-27B as a Local Reasoning Layer: 2-Week Multi-Agent Test Results
A developer replaced Claude with local Qwen3.6-27B in a multi-agent orchestrator for 2 weeks across 47 workflows. Results: strong reasoning, 12% tool-call format errors, and 12k token context limit.

Flotilla v0.5.0 Overhauls Background Execution to Beat Claude SDK Credit Caps
Flotilla v0.5.0 replaces sequential agent execution with non-blocking parallel loops, 30-minute per-agent timeouts, and local delegation to cut SDK credit usage.

MCP Slim: Local Embedding Search for MCP Tools Reduces Context Bloat
MCP Slim is a proxy that replaces full MCP tool catalogs with three meta-tools (search, describe, call), using local MiniLM embeddings for semantic search. It achieves 96% context window reduction and works offline without API keys.

Benchmarking Nemotron 3 Super 120B with 1M token context on M1 Ultra
A user tested Nemotron 3 Super 120B with a Q4_K_M quantized model using llama.cpp on an M1 Ultra, achieving a 1 million token context window that consumed approximately 90GB of VRAM. Performance benchmarks show token generation speeds ranging from 255 t/s at 512 prompt processing down to 22.37 t/s at 100,000 token context.