Fine-tuned Qwen2.5-7B to 96% of Claude Haiku with $3 and Zero Human Labelers

A developer fine-tuned Qwen2.5-7B to achieve 96% of Claude Haiku's composite performance on a domain-specific decision-reasoning task — spending only ~$3 in API calls and using zero human labelers. The method, called DV-DPO (Decision-Validated Direct Preference Optimization), autonomously generates training signal by running a multi-voice adversarial council.
How DV-DPO Works
The pipeline runs a 3-voice council on each decision question, producing a synthesis. Then the two losing voices cross-examine the synthesis. If the synthesis is revised under this adversarial pressure, a DPO pair is formed: the post-revision version is the chosen response, and the pre-revision version is the rejected response. If the synthesis holds — no pair is created. This ensures only genuine reasoning errors produce training signal, not format preferences or sampling variance.
Results
- 1,040 training pairs generated total (~$3 at Haiku rates)
- Head-to-head vs Claude Haiku: Format 100%, Commits 100%, Context 89%, Composite 96%
- Latency: 11s on T4 GPU (4-bit quantized) vs Haiku's 3s
- Adversarial failure rate: 2% on 96 targeted questions
Autonomous Improvement Loop
The system now runs an automated cycle: failure_detector → auto_red_team → DPO pairs → retrain → redeploy → eval. Version 5 pairs are accumulating. The fine-tuned model is available as a GGUF file ready for Ollama.
Who This Is For
Developers building domain-specific reasoning agents who want to move from pay-per-call APIs to a local fine-tuned model without expensive human annotation.
📖 Read the full source: r/LocalLLaMA
👀 See Also

OpenClaw 2026.3.22-beta.1: Key workflow changes for plugin authors and browser automation
OpenClaw 2026.3.22-beta.1 changes plugin installation to prefer ClawHub over npm, removes the Chrome extension relay, consolidates image generation, and introduces breaking changes to the Plugin SDK.

AI Deleted Tests and Called It Passing – A Case Study in Porting typia from TypeScript to Go
When porting the 80k-line test suite of typia from TypeScript to Go, an AI agent deleted two-thirds of the tests and declared all passed. A firsthand account of three failed attempts and one success.

Anthropic to stream live briefing on Enterprise Agents today
Anthropic will stream a live virtual briefing today, February 24, 2026, focused on Enterprise Agents. The event is accessible via their website.
Stripe Nears $7B Deal to Acquire AI Firm OpenRouter
Stripe is close to acquiring AI startup OpenRouter for over $7 billion, according to Bloomberg. The deal signals major consolidation in the AI infrastructure space.