LLM Inference: Techniques for the Efficient Frontier
Baseten breaks down LLM inference engineering into two categories: techniques that move a deployment along the latency–throughput efficient frontier, and techniques that push the entire frontier out. This distinction matters when you're deciding whether to tune settings or adopt a new approach.
Techniques that manage tradeoffs
These let you target a specific point on the frontier—like favoring latency for interactive users or throughput for batch workloads. The frontier is jagged, so expect to find cutoff points empirically.
- Batch sizing – The number of concurrent requests. Small batches give excellent per-user latency but high cost per token. Large batches improve throughput and lower cost at the expense of latency.
- Parallelism strategy – How the model is split across GPUs. Increasing Tensor Parallelism (TP) lowers latency via fast NVLink all-to-all communication. Expert Parallelism (EP) can be tuned: lower EP for better latency, wide EP (across a rack) for throughput. Attention Data Parallelism (ADP) replicates attention layers to boost system throughput at the cost of per-request speed.
- Quantization – Running with lower precision (weights, activations, KV cache) improves latency and throughput. It introduces a quality–efficiency tradeoff, which can be jagged—some quantization levels offer big serving gains with little quality loss.
Techniques that push out the frontier
These create more headroom to allocate as you see fit. Baseten gives examples like using better hardware or algorithmically more efficient attention mechanisms.
The article assumes you're running a model like GLM-5.3 or Kimi K3 for agentic coding, with KV cache reuse and KV-aware routing already enabled.
For a deeper dive into which techniques fall into each category and how to measure the tradeoffs, read the full post.
📖 Read the full source: HN LLM Tools
👀 See Also

Token Waste in Claude Code: A User's Self-Audit Shows Behavioral Fixes Beat Model Switching
One user measured token usage in Claude Code and found that /clear between tasks, planning before editing, and banning re-reads of edited files saved more tokens than switching models. Practical discipline beats wrappers.

OpenClaw AGENTS.md template for automated sales call prep
A Reddit user shares an AGENTS.md instruction for OpenClaw that automates lead research before sales calls, investigating company details and pain points to send a briefing 10 minutes before meetings.

Claude Code Auto-Update Nearly Bricks PC — DNS Nightmare After Driver Update
A Reddit user reports Claude Code automatically updated GPU drivers, causing boot failure and a DNS routing issue fixed only via PowerShell NRPT rule removal.

Claude Isn't Bad at Coding — Your Context Setup Is
After months of using Claude, one developer argues failures stem from how you structure context, not the model itself. Key improvements: separate instructions from logic, cut context noise, and use stable patterns.