Recursive Self-Improvement Framework for AI Coding Agents Using Claude Code

A developer has open-sourced a framework that enables AI coding agents to recursively improve themselves using Claude Code. The system was developed after months of research into how model providers implement recursive agent optimization.
How It Works
The framework provides a structured approach to agent improvement:
- Add tracing to your agent with 2 lines of code (or skip to step 3 if you already have traces)
- Run your agent multiple times to collect execution traces
- Run
/recursive-improvein Claude Code - The system analyzes traces, finds failure patterns, plans fixes, and presents them for approval
- Apply fixes, run agent again, and verify improvement with
/benchmarkagainst baseline - Repeat cycles to continue improvement
Autonomous Option
For fully autonomous operation (similar to Karpathy's autoresearch):
- Run
/ratchetto execute the entire improvement loop automatically - The system improves, evaluates, and keeps or reverts changes
- Only improvements survive
- Can run overnight to wake up to a better agent
Performance Results
Tested on a real-world enterprise agent benchmark (tau2) with the skill running fully on autopilot:
- 25% performance increase after a single improvement cycle
Technical Background
The original research involved building a recursive language model architecture with sandboxed REPL for trace analysis at scale, multi-agent pipelines, and other components. The developer discovered that most people building agents don't need this complexity and that Claude Code provides sufficient capability for recursive self-improvement.
The framework tells your coding agent: here are the traces, here's how to analyze them, here's how to prioritize fixes, and here's how to verify them.
Open-source repository: https://github.com/kayba-ai/recursive-improve
📖 Read the full source: r/ClaudeAI
👀 See Also

OpenAlly: Local AI Assistant for Android with Phone Control
OpenAlly is an Android app that runs an AI assistant locally on your phone via an embedded Node.js process, with 51 built-in skills and phone control capabilities through Aster companion. It connects to 19+ messaging platforms and supports 18 model providers with your own API keys.

Open-source MCP suite improves Claude Code generation quality by 15-20%
An open-source MCP suite consisting of three local servers and a prompt skill addresses the 'bad token' problem in AI code generation, with one customer reporting 15-20% quality improvement for Claude Code.

Hands-On with Tencent's Model: Strong for Agentic Workflows, Weak for Complex Coding
Tencent's model scores 8/10 for agentic tasks with low hallucination rates, but fails on complex coding like Notion API schemas. Avoid for backend logic.

bareguard: A Lightweight Safety Gate for AI Agents — Now on npm
bareguard v1.0 is a ~1000-line, single-dependency safety layer for AI agents that blocks destructive actions (rm -rf, DROP TABLE) and enforces budget limits with human escalation. Part of the bare suite, live on npm.