LLM Cost Profiler: Open-source tool tracks API spending to make case for local models

✍️ OpenClawRadar📅 Published: April 15, 2026🔗 Source
LLM Cost Profiler: Open-source tool tracks API spending to make case for local models
Ad

LLM Cost Profiler is an open-source Python tool that tracks every API call your code makes to OpenAI and Anthropic, showing exactly what you're spending, where, and why. The tool exposes which tasks are overpriced relative to their complexity, providing concrete data to make the case for local inference.

Ad

Key Features and Findings

The tool stores everything in local SQLite and is MIT licensed. According to the source, it found several specific examples of API call waste:

  • A classifier using GPT-4o that outputs one of 5 labels — a task any decent 7B local model handles easily. Cost: ~$89/week on API calls.
  • Thousands of duplicate calls to the same prompt — zero caching. Local inference with caching would make this effectively free.
  • A summarizer where 34% of calls were retries from format errors. A well-tuned local model with constrained generation eliminates this entire class of waste.

The author notes this tool gives teams concrete ammunition for investing in local inference infrastructure: "Here's the exact dollar amount we'd save by moving X task to a local model."

The tool is available on GitHub at https://github.com/BuildWithAbid/llm-cost-profiler. The author is planning to add support for tracking local model inference costs too (compute time based costing) and asked the community if this would be useful.

This type of cost profiling tool is particularly relevant for developers using AI coding agents, as it provides data-driven insights into where API spending might be inefficient compared to local alternatives.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Coordinator Server for Multi-Agent Development Prevents Overwrites
Tools

Coordinator Server for Multi-Agent Development Prevents Overwrites

A developer built a Node.js coordinator server that manages line-range locking, line shift tracking, and real-time messaging between AI agents working on the same codebase. The system prevents agents from overwriting each other's work by using HTTP-based locking with conflict detection.

OpenClawRadar
vllm-mlx fork adds tool calling and prompt cache for local AI coding agents
Tools

vllm-mlx fork adds tool calling and prompt cache for local AI coding agents

A developer has modified vllm-mlx to fix tool calling issues and add prompt caching, reducing TTFT from 28s to 0.3s for OpenClaw on Apple Silicon. The fork supports Qwen3-Coder-Next at 65 tok/s on M3 Ultra with working function calling.

OpenClawRadar
read-once: A Claude Code Hook That Prevents Redundant File Reads
Tools

read-once: A Claude Code Hook That Prevents Redundant File Reads

A developer built a PreToolUse hook called read-once that tracks files Claude Code has already read in a session, blocking re-reads of unchanged files and using diffs for changed files. The tool saves thousands of tokens per session by preventing Claude from repeatedly reading the same file content.

OpenClawRadar
cq: A Local-First Knowledge Sharing System for AI Coding Agents
Tools

cq: A Local-First Knowledge Sharing System for AI Coding Agents

Mozilla.ai's cq is an open-source tool that lets AI coding agents share 'knowledge units' about common gotchas via a local SQLite store, with optional team sharing through a Docker API. It installs as a Claude Code plugin or OpenCode MCP server.

OpenClawRadar