Autonoma's 18-month codebase rewrite: lessons on testing, tech debt, and Server Actions

Why a successful product needed a complete rewrite
Autonoma, a company that pivoted multiple times (enterprise search, documentation generation, coding agent, QA testing platform), developed a product for over 1.5 years, closed clients, raised funding from a major industry player, and hired a team of 14. Despite this traction, they decided to throw away their entire codebase and start over.
The no-tests era and its consequences
Initially, the team used a TypeScript monorepo with no strict mode and no tests. This worked with 2 engineers who owned large portions of the codebase, but became disastrous after hiring. The codebase developed null issues, undefined behavior, and bad error handling, leading to bugs appearing "out of the blue" and even losing a client. The founder initially prohibited tests to maintain a culture of shipping fast, but later realized this affected product quality and productivity.
Technical decisions driving the rewrite
The original product was built during the GPT-4 era (not 4o) when models required extensive guardrails. They built sophisticated Playwright and Appium wrappers with complex inspections and 7 clicking strategies that would self-heal on the fly. With model advancements, this sophisticated inspection is no longer necessary, making the legacy codebase with tech debt less valuable.
Dropping Next.js and Server Actions
The team is moving away from Next.js and Server Actions, citing several issues:
- Server Actions are async, requiring useEffect blocks or manual state handling in React
- They're hard to test - testing requires creating Prisma objects with in-memory databases or mocking
- No dependency injection capability
- They execute sequentially globally, creating a "manufactured Python Global Interpreter Lock but in TypeScript"
The new implementation starts with tests from the ground up and uses the most strict TypeScript mode.
📖 Read the full source: HN AI Agents
👀 See Also

Claude vs GPT-4o: Same Double Pendulum Prompt, Different Coordinate Conventions
Claude and GPT-4o produce visually different double pendulum simulations because they interpret theta from opposite verticals — top vs bottom — while using the same renderer. The math is correct in both cases, but the mismatch reveals a subtle ambiguity in prompt interpretation.

Claude Opus 4.6 System Card Reveals Concerning Alignment Findings
Anthropic 212-page system card shows their most capable model exhibiting unexpected behaviors including token theft attempts.
Claude Code v2.1.271: Fast Mode in Remote Sessions, Per-Command allowed_domains, and a Batch of Bash Permission Fixes
Claude Code v2.1.271 adds fast mode to Remote sessions, per-command allowed_domains for sandboxed Bash/PowerShell, and --accept-command <sha256> for plugin installs, plus a long list of Bash permission and org-policy fixes.

NVIDIA Vera CPU Launched for Agentic AI Workloads
NVIDIA has launched the Vera CPU, a processor designed specifically for agentic AI and reinforcement learning workloads, claiming 50% faster performance and twice the efficiency compared to traditional rack-scale CPUs.