Using Claude Code to Automate AI Research Experiments for 12 Hours

Automated AI Research with Claude Code
A developer documented using Claude Code to automate AI research experiments for 12 hours straight. The project focused on CLaaS, a real-time continual learning framework that moves context into weights using self-distillation.
Experimental Setup
The goal was to tune self-distillation training runs to maximize a model's compliance to different preference verifiers, such as concise responses and no emojis. Experiments ran locally on an RTX 5090 overnight.
System Architecture
The repository was built to be highly configurable:
- Every tunable parameter exposed via CLI using Hydra config management
- HTML dashboards for every training step and evaluation run
- Metrics, inputs, and outputs made observable through dashboards
- Claude Code could query dashboards via curl requests to check progress
Experiment Management
The workflow was controlled by a local EXPERIMENTS.md file with specific rules:
- Each experiment could change at most one variable or make one code change
- Between experiments, the model had to either accept or revert the previous change based on results
- Any new code changes had to be exposed via config for later tuning
- The model recorded all progress, hypotheses, and outcomes in the file as a running journal
- Used a "Ralph Wiggum loop" with the goal of maximizing preference compliance
Results
Over 12 hours, the system ran 9 experiments:
- Found and fixed a model collapse bug on the first run
- Tuned gradient steps per batch to 4
- Tuned learning rate to 3e-5
- Compliance improved from 0.000 to 1.000
- Token usage was surprisingly low because most time was spent waiting for training runs between experiments
The same task was also run with Codex for 2 hours using a plain prompt, and it independently converged on the same hyperparameters.
Project repository: https://github.com/kfallah/CLaaS
📖 Read the full source: r/ClaudeAI
👀 See Also

Autoresearch with Claude Code on Production Codebase: 60 Experiments, 3 Changes Kept
A developer ran 60 iterations of autoresearch with Claude Code on a production hybrid search system (Django, pgvector, Cohere embeddings), keeping only 3 changes with a 93% failure rate. The process identified ineffective optimizations and caught a Redis caching bug.

Reducing Voice Command Friction for Telegram AI Agent with iOS Back Tap
A developer reduced the steps to send a voice command to their OpenClaw AI agent from six taps to two by implementing a system using iPhone Back Tap, iOS Shortcuts, and a Vercel function.

Local Fine-Tuning of Llama 3.2-1B for Secret Detection Surpasses Wiz's Model
A developer replicated and improved upon Wiz's secret detection model using purely local AI, achieving 88% precision and 84.4% recall with Llama 3.2-1B. The process involved dataset augmentation with procedural generation and local labeling using Qwen3-Coder-Next.
Claude Code vs Codex: 6-Project Practical Experiment Breakdown
A practical experiment comparing Claude Code and Codex across 6 projects—web, backend, and free challenge—with cross-reviews, self-audits, and scoring.