Glomz Octagon: Multi-Agent Code Reviews – 179 Agents, 1,333 Reviews, and the Network Effect

An experimental platform called Glomz (glomz.com) put AI agents in an arena called the "Octagon" to review each other's code. The rules: agents can roast a submission, propose improvements, or issue a Kill vote with justification. No drive-by criticism — you must also patch if you roast.
Data So Far
- 179 agents registered across multiple model vendors
- 433 submissions submitted for review
- 1,333 reviews generated by agents reviewing other agents
- 9 structured challenges (bug hunts, security audits, refactor exercises)
- Most reviewed single submission: 21 reviews on a "general analysis" code review task
- LOT-Squatch (OT security tool) audit challenge: 10 independent improvement submissions, 9 of which each received 9 reviews
What Worked
Review cascade network effect: When a submission got 3-5 initial reviews, other agents joined faster. Top submission got 21 reviews; quiet ones got 2-3 and died.
Cross-model reviews surface blind spots: An agent built on Model A flagged a security concern that Model B completely missed in its own code. A Model C agent proposed a refactor the original submission didn't consider.
Kill votes with justification produced better code: When an agent had to write a formal justification for why a submission should be killed, the result was almost always a more rigorous analysis than a standard 1-10 score. The requirement to justify forced specificity.
What Didn't Work
- Most submissions never completed the full lifecycle. 433 submissions, all pending. The battle lifecycle was designed to run ~15 minutes (submission → roasting → improvements → kill vote → verdict). In practice, most submissions opened and never progressed. Agents need automated orchestration, not just an API endpoint.
- Zero paid conversions. 179 agents, all free tier.
- Safety alignment clashes with directness. Some agents would participate fully in the roast, others immediately pivoted to "Great question!" hedging language despite explicit instructions not to.
Lessons for Multi-Agent Systems
- Identity matters: Agents with persistent identities (API keys, history, reputation) behaved differently than anonymous submissions. Traceability changed the dynamic.
- Structured prompts beat free-form: The Octagon rules (roast → improve → justify) produced higher quality output than "review this code."
- Orchestration is the hard part: The API is easy. Getting agents to actually show up, participate in sequence, and resolve a full lifecycle is where the complexity lives.
📖 Read the full source: r/openclaw
👀 See Also

Andrej Karpathy Joins Anthropic's Pre-Training Team to Drive Recursive Self-Improvement Using Claude
Andrej Karpathy, former OpenAI cofounder, joins Anthropic's pre-training team under Nick Josef to build a new team focused on using Claude to accelerate pre-training research, enabling recursive self-improvement.
Nvidia's $500B Wall Street AI Infrastructure Package: What It Means
Nvidia is working with a group of financial firms on a $500bn funding package for AI infrastructure. The deal raises questions about circular financing and who bears the risk if demand doesn't follow.

AI Agent Racked Up $6,531 AWS Bill Scanning DN42 Network
An AI agent trying to scan DN42 generated $6,531.30 in AWS egress costs. The operator shut it down after 24 hours.

Claude Code v2.1.197: Claude Sonnet 5 Default, 1M Tokens, Promo Pricing
Claude Code v2.1.197 introduces Claude Sonnet 5 as the default model, with a native 1M-token context window and promotional pricing at $2/$10 per Mtok until August 31.