OpenEvol: Offline Self-Improvement Pipeline for LLMs Using Conversation History

What OpenEvol Does
OpenEvol is an offline self-improvement pipeline for large language models that automatically converts AI conversation history into training data. The tool mines high-value exchanges from conversations, judges their quality, and generates fine-tuning datasets without manual labeling or proprietary data flywheels.
How It Works
The pipeline runs through four automated stages:
- Mine high-value exchanges from conversations
- Judge quality using rules with an optional teacher LLM
- Synthesize SFT, preference, and pretraining datasets
- Fine-tune with one command
This creates a closed loop where the model learns from its own experience.
Technical Details
No GPU is required to get started - the full pipeline runs on CPU with a mock or OpenAI-compatible teacher backend. You can bring a GPU when ready to train.
Five teacher backends are supported:
- Mock
- Rule-based
- OpenAI-compatible API (any local proxy works)
- HuggingFace Transformers
- vLLM
Usage Options
Three ways to use OpenEvol:
- CLI for offline batch runs
- REST API server for automation
- OpenClaw desktop plugin that lets you trigger pipeline runs directly from chat
Quality Control
Every batch is automatically scored. If the approval rate drops below 80%, training is blocked and flagged for human review, giving users control over what data gets used for training.
This type of tool is useful for developers who want to improve their AI coding agents using actual conversation history without sending data to external services.
📖 Read the full source: r/openclaw
👀 See Also

Setting Up OpenClaw as an Always-On AI Assistant
OpenClaw, configured as an always-on AI assistant for a small dev team, is set up on a Railway server with Claude as the backend and integrates with Google Workspace, GitHub, and more.

Yavio: Open-Source Product Analytics SDK for MCP Apps
Yavio is an open-source product analytics SDK for MCP and MCP Apps that automatically captures tool calls, errors, and resource reads with one function call. The MIT-licensed project provides a dashboard with per-tool breakdowns, funnels, retention, and error tracking.

Claude Code Used to Simulate 4,000+ Blind Werewolf Games with LLMs
A developer used Claude Code to build a simulator where LLMs play blind one-night Werewolf, running ~4,600 games across OpenAI and xAI models. The experiment revealed consistent name-based voting patterns despite minimal game signals.

Buyer Eval: Claude skill for B2B vendor evaluation using AI agent conversations
A Claude skill that evaluates B2B software vendors by researching your company, asking domain-specific questions, and directly interrogating vendor AI agents through the Salespeak Frontdoor API. It cross-references claims against independent sources and produces evidence-based scorecards with transparent verification levels.