Claude Code Rewrites PostHog's SQL Parser for 70x Speedup – How Property-Based Testing and Parallel Agents Worked

PostHog engineer Robbie Coomber used multiple long-running Claude Code sessions in parallel to rewrite their SQL parser, achieving a ~70x speedup over the existing ANTLR-generated C++ parser. The new parser is 16K lines of hand-rolled Rust code, plus 5K lines of tooling and tests.
Why rewrite?
PostHog transpiles user SQL to raw ClickHouse SQL for logical data views and optimizations. The parser turns SQL into an AST for downstream access control and optimization. The old ANTLR-generated parser used a generic graph-walking interpreter (an ATN — NFA-with-a-stack) with arbitrary dynamic lookahead, which was slow despite being in C++. Hand-rolled recursive-descent parsers are inherently faster.
Approach: parallel agent sessions + oracle TDD
- Tested two approaches in parallel: one focused on performance (recursive-descent with Pratt expression parsing), the other on correctness (mimicking ANTLR's behavior with explicit code). Both worked equally well.
- Used the existing C++ parser as an oracle to generate disagreements — find SQL where parsers differed, fix the new parser, repeat.
- Property-based testing generated countless SQL variations, including a test for
SELECT SELECT FROM FROM WHERE WHERE AND AND(valid SQL). - The new parser agrees with the oracle for all realistic queries; differences occur only for pathological queries.
Results
70x speedup in parsing. The final parser is a recursive-descent design with lookahead and backtracking only where necessary. Coomber notes that without AI, writing and maintaining such a hand-rolled parser would take months and likely not be worth the effort. With Claude Code, it became practical.
Key takeaway
This case study shows that parallel AI coding sessions, combined with property-based testing and an oracle for test-driven development, can dramatically improve performance-critical code. The technique — using agents to rewrite core infrastructure while relying on automated disagreement detection — is reusable for other projects.
📖 Read the full source: HN AI Agents
👀 See Also

ScreenMind: Local-First AI Memory That Indexes Your Entire Computer Activity
ScreenMind captures your screen, meetings, and voice notes using Gemma 4 E2B locally via llama.cpp. Runs on 4GB+ VRAM with Q4 quantization. Search past activity, chat with history, and connect to Claude/Cursor via MCP.

Merlin: Local-first LLM context dedup – measure up to 71% chunk overlap, free & open-core
Merlin is a local-first context dedup tool that measured 22-71% chunk overlap across 22M passages from real agent/RAG sessions. Ships as HTTP proxy (Ollama/vLLM/SGLang/llama.cpp), MCP server (Claude/Cursor/OpenClaw), or standalone CLI. MIT open-core with daily usage caps.

Cortex v1.2 adds LLM enrichment, Q&A with citations, and conflict resolution
Cortex, a local memory layer for OpenClaw agents, has released v1.2 with LLM-augmented enrichment by default, a question-answering command with citations, and improved deduplication and conflict resolution. The tool now includes unified configuration management and intent-based search pre-filtering.

Auto-co: A 50-Line Bash Script That Turns Claude Code Into Autonomous AI Companies
Auto-co is a 50-line bash script that wraps the Claude Code CLI in a loop, allowing it to run autonomously with 14 AI agents playing roles like CEO, engineer, and critic. It has built four products from scratch, including FormReply and Changelog.dev, at a total cost of $268 across 270+ cycles.