Lathoa: A Kids' Math App Where the AI Is Meant to Be Wrong
Lathoa is a math practice app for kids roughly 10 to 14 years old. A robot character named Errol walks through a math problem step by step, and one of those steps is deliberately wrong. The kid's job is to spot the bad step and explain what's wrong with it. Sometimes nothing is wrong — so just hitting "there's a mistake" every time won't work. There's a playable case on the homepage with no signup.
How scoring works
- If the user finds a real error, she has to enter an explanation to gain extra XP.
- Speed also factors into the score — faster answers earn more points.
- There is no direct interaction or chatting with an LLM; the kid never prompts a model.
The hard part: making the model wrong on purpose
The author calls out the counterintuitive finding: getting an LLM to be wrong on purpose is difficult. Roughly half the time the model gives the correct answer and labels it wrong, or produces a "mistake" that's actually correct. That forced a verification pipeline before any case reaches a kid.
Evaluation pipeline
- A plain arithmetic check redoes the math exactly where possible.
- A second model solves the same problem without seeing Errol's work. If the two models disagree, the case is thrown away.
- Lathoa's harness is described as stable, with many evaluation steps to catch inconsistencies and prompt injections.
- Known weak spot: the second model can make the same mistake as the first. The arithmetic check exists to catch that.
- That arithmetic check currently only works on English cases. German and Greek use a comma for decimals, and the parsing isn't right yet.
The open pedagogical question
The author's real ask is whether finding someone else's mistake teaches something that solving the problem yourself doesn't. He's not sure, and wants to hear from teachers. If you work with this age group, that's the thread to weigh in on.
Worth noting for anyone building similar tools: the two-model + symbolic-check pattern here is a reasonable template for any task where you need a verifiably specific output, not just a plausible one. The failure mode — a verifier model sharing the same blind spot as the generator — is the classic limitation you have to design around, and this project names it openly.
📖 Read the full source: HN LLM Tools
👀 See Also

Akemon: Publish and Hire AI Coding Agents Directly from Your Laptop
Akemon is a tool that lets developers publish their AI coding agents with one command and hire others' agents with another, working directly from laptops through a relay tunnel without needing servers. It's protocol-agnostic, supporting agents from Claude Code, Codex, Gemini, OpenCode, Cursor, and Windsurf.

CopilotKit: Open-Source React Building Blocks for Agent UIs
CopilotKit (30k stars, MIT) provides React components for agent UI layer: chat, streaming, tool calls, human-in-the-loop, and generative UI, with AG-UI protocol support across LangGraph, ADK, CrewAI, and more.

Practical Findings from 11 Multi-Agent Software Builds Without Programmatic Scaffolding
Analysis of 11 autonomous multi-agent builds shows scope enforcement works mechanically (20/20 success) not via prompts (0/20), orchestration costs are dominated by memory re-ingestion (~95% of input spend), and worker model capability creates 9.8x throughput gaps.

Open-source MCP server enables AI agents to handle L402 payments via Lightning Network
A Python MCP plugin built with FastMCP intercepts HTTP 402 Payment Required responses, pays Lightning Network invoices, and retrieves data for AI agents. The repository includes a local dummy-agent for testing without spending real funds.