A Misalignment of AI in Mathematics: Tao, the Economist, and a 622-Comment HN Debate
On September 11, 2026, Terry Tao's blog and The Economist both ran pieces on the same topic: AI misalignment in mathematics. The Hacker News thread aggregating them hit 560 points and 622 comments — the discussion is doing most of the work here, since the original posts are paywalled upstream and the HN comments carry the technical substance.
What's actually being argued
The framing is about misalignment, not capability. The concern isn't that LLMs are bad at math — it's that the incentives around AI-assisted mathematics point somewhere other than mathematics. Concretely, the kinds of claims in circulation in this debate are:
- Formalization pressure: Lean 4 and Mathlib have made machine-checkable proofs viable, which changes what "done" means. An LLM that produces a plausible-looking proof sketch is not the same as one that produces a proof that compiles.
- Benchmark displacement: Tools like AlphaProof and AlphaGeometry score on competition problems (IMO-style). That's a narrow, well-specified target. Research mathematics is not.
- Review bottleneck: If AI accelerates proof generation faster than humans can verify it, you get a growing backlog of unverified claims — the opposite of what formalization is supposed to buy you.
- Credit and attribution: Tao's own framing tends to be about where human judgment still has to sit in the loop, and what happens when it doesn't.
Why developers should care
The pattern generalizes. If you've shipped an agent that writes code, you already know the failure mode: the model produces something syntactically valid and semantically wrong, and the cost shifts from writing to reviewing. Mathematics is the cleanest test case because it has a verifier — Lean, Coq, Isabelle — that doesn't care about vibes. If misalignment shows up even there, it shows up everywhere else first.
The HN thread is the useful artifact. 622 comments on a paywalled pair of posts usually means the comments are where the actual arguments live. Read it.
📖 Read the full source: HN LLM Tools
👀 See Also

Gemma 4 Chat Template Bug: Tool Parameters with anyOf/null Rendered as Empty type
A bug in Gemma 4's chat template drops $ref, anyOf, and $defs from tool parameter schemas, rendering nullable refs as empty type fields. A Jinja fix restores correct schema parsing for all inference engines.

ACP Bug Investigation: Protocol Mismatch Causes 'metadata is missing' Error with Local Ollama
A confirmed bug in the ACP/OpenClaw integration prevents acpx spawn commands from working with local Ollama models due to a protocol mismatch where acpx expects JSON but receives text output.

AI Data Center Water Use in California: Estimates from Physics and AI Models
A California WaterBlog analysis using physics and four AI models estimates AI data center water use in California at 2,300–400,000 acre-ft/year, with a realistic range of 32,000–290,000 acre-ft/year — modest compared to agriculture.

Claude Code allegedly refuses requests or charges extra when commits mention 'OpenClaw'
A tweet by Theo claims Claude Code either refuses requests or charges extra if your git commits mention 'OpenClaw', sparking discussion on HN.