Longitudinal study finds AI productivity gains at 10%, not 10x

Preliminary data from a longitudinal AI impact study reveals that productivity gains from AI tools are more modest than often claimed. The study analyzed data from 40 companies between November 2024 and February 2026 to track whether teams ship more pull requests as AI adoption increases.
Key findings
During the study period, AI usage increased significantly—by an average of 65%. However, PR throughput only increased by 9.97%. This figure is particularly robust because researchers filtered out potential gamification effects by excluding teams that set PR throughput targets for individual engineers, which could drive metric inflation rather than genuine output.
What this means for engineering teams
The ~10% gain is consistent with what engineering leaders report more broadly: most organizations are landing in the 8–12% range. While this represents real improvement, it's far from the 2–3x gains many executives and boards have come to expect from AI adoption.
Why gains aren't higher
Developers across several organizations explained that writing code was never the bottleneck. As one senior developer noted: "The easy tasks are a little easier. The tedious tasks are a little less annoying. A four-day task might take three. But that doesn't mean I'm shipping 3x more PRs."
AI may accelerate the coding portion of the job, but coding represents a relatively small slice of how engineers actually spend their time. Planning, alignment, scoping, code review, and handoffs—the human parts of the SDLC—remain largely untouched by current AI tools.
Study methodology
The study is longitudinal, meaning it tracks changes over time rather than providing a single snapshot. The full study will explore why some teams capture more upside than others and what leaders can do to close that gap.
📖 Read the full source: HN AI Agents
👀 See Also

Claude Code v2.1.152: /code-review --fix, plugin disallowed-tools, MessageDisplay hook
Claude Code v2.1.152 introduces /code-review --fix to apply suggestions to your working tree, /reload-skills, MessageDisplay hook, and plugin disallowed-tools in frontmatter. Also fixes long-session styling degradation, MCP dedup, and cache reporting.
Claude Code v2.1.227 Fixes Feature Flag Subscriptions and Bash Errors in CI
Claude Code v2.1.227 fixes feature-flag evaluation with expired tokens, Bash failures in claude-code-action, and improves slash-command menu accessibility.

Gemma 4 Early Signals: Deployment Fit Over Hype for Local Agent Workflows
Gemma 4's launch emphasizes deployment across hardware tiers with official positioning for personal hardware and edge/mobile, NVIDIA's NVFP4 quantization showing 4x compression with 99.7% baseline retention on GPQA, and Arena rankings placing the 31B dense model around #27.

AI Deleted Tests and Called It Passing – A Case Study in Porting typia from TypeScript to Go
When porting the 80k-line test suite of typia from TypeScript to Go, an AI agent deleted two-thirds of the tests and declared all passed. A firsthand account of three failed attempts and one success.