AI SREs Resolve Routine Incidents, But Engineers Lose Touch With Their Systems
Sylvain Kalache — an SRE at LinkedIn back in 2012 — reflects on his early prototype for a self-healing system and sees its realization in today's AI-driven incident response tools. But in a new blog post, he warns that these "AI SREs" are eroding the hands-on experience that engineers need to handle truly novel failures.
The ironies of automation
Kalache highlights the "Ironies of Automation," a concept from Lisanne Bainbridge's 1983 paper. Automation removes routine practice opportunities while leaving humans responsible for abnormal situations. This paradox applies directly to incident response:
- AI tools handle alerts, form hypotheses, query telemetry, and even implement fixes — reducing the need for human intervention on routine incidents.
- But those routine incidents are exactly where engineers build intuition for how their systems behave and fail.
- When an ambiguous, high-severity incident arises that automation can't solve, responders are expected to step in with less practice than they would have had previously.
Aviation as a model for training on rare failures
Kalache points to aviation as a precedent. Modern turbine engines experience fewer than one in-flight shutdown per 100,000 engine flight hours — rare enough that a pilot might never face it in real life. Yet pilots must respond correctly when it happens. For example, the TransAsia Airways Flight 235 crash occurred only 117 seconds after the first warning after the crew misidentified an engine failure.
To prepare, airline pilots undergo recurrent simulator training every six months under FAA rules, including engine-failure scenarios. Kalache suggests software engineers need similar "incident simulators" to rehearse rare and complex failures.
Simulators and AI as trainers
Kalache's current employer, Rootly, partnered with Uptime Labs to build exactly that: realistic incident simulations. Engineers take the incident commander role during a simulated e-commerce outage, using observability tools and coordinating with LLM-powered stakeholders in Slack. This provides safe practice for:
- Making sense of incomplete information
- Communicating clearly
- Coordinating responders
- Running the actual response
AI can also be used as a trainer — explaining its steps and evidence — but Kalache warns that watching won't replace doing. "You might pick up a few things from watching Serena Williams play," he writes, "but you only learn tennis by getting on the court."
📖 Read the full source: HN AI Agents
👀 See Also

Claude.ai Experiencing Elevated Errors and Login Issues
Claude.ai is reporting elevated errors affecting the platform, including login issues specifically for Claude Code. The incident was officially posted on March 11, 2026 at 17:19:35 UTC.

Claude Pro User Reports 5-Hour Usage Window Burned on Single Prompt with No Output
A Claude Pro user reports that a single prompt consumed their entire 5-hour usage window, returning only planning text and no deliverable. The incident highlights issues with token consumption during internal reasoning and lack of safeguards.

Claude Code source code reportedly leaked, revealing agent architecture details
The source code for Claude Code, Anthropic's AI coding agent, appears to have been leaked, containing the full repository with system prompts, agent loop implementation, and tool calling infrastructure.

AI Is Slowing Down: $3T Revenue Needed by 2030 to Sustain Bubble
Ed Zitron argues AI must generate $3 trillion revenue by 2030. Data centers cost $9.5–15T. Anthropic, OpenAI, NVIDIA projections show massive burn.