Homa: Replacing TCP for AI Cluster Networking
Homa is a transport protocol designed to replace TCP in datacenter and AI cluster networks, where microsecond-scale tail latency matters more than byte-stream semantics. This video walks through the design and its motivation.
What Homa changes vs TCP
TCP was built for the wide-area internet: it optimizes for bandwidth and fairness across long-lived flows. Homa is receiver-driven. Instead of senders probing for capacity, the receiver grants credit back to senders and schedules incoming messages using SRPT (shortest-remaining-processing-time first). Short messages — the common case for RPC and collective ops in AI training — get priority over large bulk transfers.
Why this matters for AI clusters
AI training and inference traffic is bursty and message-oriented. AllReduce, parameter server updates, and KV-cache transfers all look like many small messages rather than a few long byte streams. TCP's per-flow congestion control and in-order byte delivery add head-of-line blocking and queueing delay that hurt job completion time at scale.
The original USENIX ATC '21 paper by Ousterhout et al. is the core reference. It reports order-of-magnitude reductions in 99th-percentile message latency compared to TCP-based stacks under datacenter workloads, plus higher throughput when many short messages share the fabric.
Where it stands now
The linked LWN article covers the ongoing effort to bring Homa-like scheduling into the Linux networking stack — the long-running discussion about whether to extend TCP, add a new transport, or build it as a kernel module. The Register piece covers the Stanford project angle.
If you're building or operating GPU clusters and your job completion times are dominated by network tail latency rather than compute, this is worth following. The paper is the fastest way to understand the design; the LWN thread shows what kernel integration would actually look like.
📖 Read the full source: HN LLM Tools
👀 See Also

Claude AI Spends 81 Minutes on 'Real Thinking' – User Report Spikes Around Major Updates
A user reports Claude AI spent 1 hour 21 minutes on a simple task, speculating that performance spikes happen briefly after major updates. Example: a research request scanned 5,113 sources in one session but later only 100-200 sources for similar queries.

EU Forces Google to Share Search Data, Open Android to Third-Party AI
New DMA specification measures require Google to share search data with competitors and open Android to third-party AI assistants with full system integration by mid-2027.

Claude Service Incident: Elevated Errors Across Platforms
Claude experienced elevated errors across claude.ai, console, and Claude Code platforms on March 2, 2026, with issues affecting login/logout paths and some API methods. The incident was resolved after approximately 4 hours.

Nvidia's Nemotron 3 Super: 120B Parameter Model with 12B Active Inference
Nvidia's Nemotron 3 Super has 120 billion total parameters but only activates 12 billion during inference, achieving 120B model knowledge at roughly 12B compute cost through efficient routing rather than compression.