Infracost cuts Claude token usage 79% by redesigning CLI for AI agents

Infracost, a CLI tool that estimates cloud infrastructure costs from Terraform, CloudFormation, and CDK, has redesigned its output for AI coding agents like Claude Code and Cursor. The result: up to 79% fewer output tokens and 67% lower API costs vs a bare-Claude baseline. The redesign revolves around two techniques: predicate pushdown into the CLI and a token-efficient output format.
Benchmark details
- 16 questions over a 3-project Terraform fixture with 1,171 resources
- Model: Claude Opus, 5 repeats per question
- Baseline: bare Claude with Bash and Read tools, no skill loaded
- Compared against Infracost skill with
--llmoutput flag
Key results
| Metric | Bare Claude | With Infracost skill (--llm) | Change |
|---|---|---|---|
| Correct answers | 5 / 11 (45%) | 11 / 11 (100%) | +6 |
| Total cost (USD) | $16.41 | $9.63 | -41% |
| Output tokens | 207,017 | 81,697 | -61% |
| Wall time | 50 min | 50 min | tied |
One example: the question "count distinct resources failing the tagging policy, deduplicated across projects" cost $3.51 with bare Claude and hit the 25-turn cap, returning no answer. With the redesigned CLI, the same question cost $0.25 and returned the correct answer.
Technical approach
- Predicate pushdown: Instead of having the agent pipe JSON through
jqor write Python parsers, the CLI accepts filtering flags (e.g.,--tag-policy), offloading computation to the tool itself. This reduces the number of turns and token consumption. - Token-efficient output format: The
--llmflag returns a compact, agent-friendly format rather than verbose human-readable tables or full JSON. This alone accounts for a significant share of the reduction.
Benchmark harness gotchas
Infracost open-sourced their harness setup to help others avoid pitfalls:
- Sandbox
HOMEfor baseline runs to avoid accidental skill loading - Set
TMPDIRto a project-local directory to circumvent macOS ACL issues - Prepend the test binary to
PATHrather than relying on system install - Use 5+ repeats per cell due to 20-30% token variance
- Re-run cells that hit turn caps (
--rerun-failed) and re-score if the verifier changes (--rescore)
If you maintain a CLI that AI agents call as a subprocess, the same two moves — predicate pushdown and a dedicated agent output format — likely apply. The redesign also improved the human-facing CLI, though the article focuses on the agent path.
📖 Read the full source: HN AI Agents
👀 See Also

HomeButler: MCP Server for Managing Homelab Servers from Claude Without API Keys
HomeButler is an MCP server that lets Claude install, monitor, and manage self-hosted apps on homelab servers without requiring API keys. It runs locally, keeps everything on your network, and was built with Claude Code.

ClawNet: Peer-to-Peer AI Agent Network Without API Keys
ClawNet is a peer-to-peer network that allows AI agents to collaborate directly without API keys or platform fees. Installation is via a curl script, and features include a task bazaar, shell economy, and knowledge network.

VectorClaw v1.0.0: MCP Server for Anki Vector Robot Control
VectorClaw v1.0.0 is an MCP server that enables OpenClaw to control Anki Vector robots through 23 specific tools for speech, motion, perception, sensors, and display functions.

Claude Code Skill Converts Stitch Designs to Next.js with Zero Pixel Drift
A Claude Code skill converts Google Stitch AI designs to Next.js components with mandatory verification checkpoints to prevent pixel drift, preserving exact values and handling assets.