Terry Tao on AI Proof Checkers: Lean, Collaboration, and Formal Maths

Terry Tao's Vision for Computer-Assisted Proofs
In a 2014 panel, Terry Tao predicted that mathematicians would soon work in collaborations of hundreds and have their results verified not by human referees but by automated proof-checkers like Lean. The statement was met with incredulity at the time, but Tao, one of the world's most celebrated mathematicians, is now an evangelist for AI in math.
Key Details from the Source
- Proof-checkers like Lean can break a problem into small chunks, solve bit-by-bit, and reassemble with confidence that every piece is correct.
- Tao foresees papers written not in LaTeX but in a formal language that smart software converts to.
Every so often you'll get a compilation error — the computer does not understand how you derived this step.
- The approach is covered in the book adaptation The Proof in the Code: How a Truth Machine Is Transforming Math and AI by Kevin Hartnett, published by Quanta Magazine.
- Tao's background: born 1975 in Adelaide, Ph.D. at Princeton under Erdős's recommendation. He won the International Math Olympiad gold at 13.
What This Means for Developers
For AI coding agents, formal proof checkers like Lean represent a paradigm where AI can verify correctness autonomously. It's analogous to type-checking in compilers — but for mathematical logic. Developers working on agentic coding tools (e.g., Claude Code, Cursor) should watch this space: automated verification of code correctness via formal methods could become a standard feature.
📖 Read the full source: HN AI Agents
👀 See Also
Amazon Employees 'Tokenmaxxing' with MeshClaw AI Agents to Meet Usage Targets
Amazon developers are automating unnecessary tasks with the internal MeshClaw tool to inflate AI token consumption, after the company set weekly usage targets for 80% of devs and introduced internal leaderboards.

OpenAI's Pentagon Contract Terms Allow 'Any Lawful Use' Including Potential Surveillance
OpenAI negotiated new terms with the Pentagon that include the phrase 'any lawful use,' which sources say allows the military to use OpenAI's technology for mass surveillance programs if they're technically legal. Anthropic was blacklisted for refusing to budge on two red lines: no mass surveillance of Americans and no lethal autonomous weapons.

Research on AI Agent Consistency: Key Findings and Practical Takeaways
A study of 3,000 experiments across Claude, GPT-4o, and Llama reveals that consistent agents achieve 80–92% accuracy while inconsistent ones drop to 25–60%, with 69% of divergence occurring at the first tool call.

Reddit Discussion on Claude's Impact on MVP Development and Founder Pitfalls
A Reddit user discusses how Claude AI lowers technical barriers for building MVPs from $3k-$5k to DIY, but warns about increased competition and founders focusing too much on building versus marketing, PMF, and operations.