Sakana AI Launches RSI Lab: Recursive Self-Improvement with Foundation Models

Sakana AI has formally established its Recursive Self-Improvement (RSI) Lab, a dedicated research group tasked with redesigning the AI development process itself using AI. Rather than brute-forcing monolithic models, the lab builds open-ended, adaptive architectures that collectively self-improve — drawing on a lineage of published milestones.
Key Research Milestones Backing RSI
- LLM-Squared (2024): Developed with Oxford and Cambridge, this framework lets LLMs invent better ways to train LLMs (LLM²). It produced DiscoPOP, a preference optimization algorithm discovered and written entirely by an LLM through a generational evolutionary loop.
- Darwin Gödel Machine (2025): In collaboration with UBC, DGM maintains an evolving lineage of agent variants that autonomously rewrite their own codebase. On SWE-bench, it more than doubled baseline performance — a 30 percentage point absolute improvement.
- ShinkaEvolve (2025): Open-source framework demonstrating sample-efficient program evolution. Solved complex optimization problems using only 150 samples and generated a novel load-balancing loss function improving Mixture-of-Experts (MoE) models.
- ALE-Agent (2025): Optimization agent that secured 1st place out of 804 human participants in AtCoder Heuristic Contest 058. It leverages massive inference-time scaling and self-learning from trial-and-error failures to autonomously derive novel algorithms.
- Digital Red Queen (2026): Collaboration with MIT establishing open-ended adversarial coevolution in Core War. LLMs author competing code, driving emergent complex software strategies and convergent evolution — foundational for cybersecurity RSI.
- The AI Scientist (2024–2026): Fully automated open-ended scientific discovery, from idea generation, experiment execution, full paper writing, to peer review.
Why This Matters for Developers
RSI represents a shift from static, human-led R&D to autonomous self-improving intelligence engines. The lab's approach — evolutionary optimization loops, self-rewriting agents, and automated science — directly impacts how AI coding agents are built and improved. Rather than waiting for manual tuning, these systems continuously refine their own architectures.
📖 Read the full source: HN AI Agents
👀 See Also

KV Cache Architecture Evolution: From GPT-2 to Mamba
Analysis of KV cache memory costs shows GPT-2 used 300 KiB/token, Llama 3 reduced it to 128 KiB/token with grouped-query attention, and DeepSeek V3 achieved 68.6 KiB/token with multi-head latent attention. Mamba/SSMs eliminate KV cache entirely with fixed-size hidden states.

MiniMax M2.7 Model Shows Strong Performance as AI Coding Agent
A developer tested MiniMax M2.7 as their main AI coding agent and found it outperformed GPT 5.4 and Gemini 3.1 Pro in speed and tooling tasks, with benchmark scores of 56.22% on SWE-Pro and 57.0% on Terminal Bench 2.

Nine Common AI Coding Agent Failure Patterns and Pre-Execution Validation
A Reddit post identifies nine specific failure patterns that commonly cause AI coding agents to fail, including incomplete enum handling, silent null paths, and hallucinated imports. The author reports implementing a validation pass before execution catches about 70% of these failures.

Claude Code v2.1.161: OTEL Attributes, Parallel Tool Fixes, and MCP Secret Redaction
v2.1.161 includes OTEL resource attributes as metric labels, independent parallel tool results, MCP secret redaction, and multiple bug fixes for subagents, Windows hooks, and OpenTelemetry log events.