How to Cut OpenClaw Agent Costs by 80% with Model Switching

A Reddit user spent two weeks manually logging every OpenClaw agent interaction to figure out where their money was going. The results are a clear blueprint for optimizing spend on AI agents.
The Breakdown
Over 14 days on a Telegram + Discord agent, token usage broke down as follows:
- Heartbeats (30-min polls) — 38% of usage. Running on Opus at ~$6.75/M tokens. Complete waste for a status ping.
- File reads and summaries — 29% of usage. Also on Opus. Flash handles these identically.
- Actual conversations — 22% of usage. Here model quality matters.
- Complex tasks — 11% of usage. Where Opus genuinely outperforms Flash.
In total, 67% of spend went to tasks where DeepSeek V4 Flash ($0.14/M) would deliver identical quality to Opus ($6.75/M effective after tokenizer).
The Fix: Default to Flash, Escalate Only When Needed
Set your primary model to deepseek/deepseek-v4-flash in openclaw.json:
"agents": {
"defaults": {
"model": {
"primary": "deepseek/deepseek-v4-flash"
}
}
}Then use /model anthropic/claude-opus-4-7 mid-session when you hit something truly hard. The switch is instant — no restart, same session. Type /model deepseek/deepseek-v4-flash when you're done to drop back to cheap.
Results
Costs dropped from ~$170/month to ~$35/month. The quality difference on heartbeats, file reads, and simple questions was literally zero.
The user notes that BetterClaw's free tier (with BYOK) now shows per-task API spend, which would have caught the heartbeat waste immediately. But the core move — switching primary to Flash and /model-ing up to Opus only when needed — is the real takeaway.
📖 Read the full source: r/openclaw
👀 See Also

MTP Acceptance Rate: 50% Threshold Determines Speculative Decoding Benefit
MTP (Multi-Token Prediction) via speculative decoding on Gemma-4 26B shows benefit only when draft token acceptance rate exceeds 50% — based on mlx-vlm benchmarks on M4 Max Studio.

Hide OpenClaw Exec Lines in Telegram Chat: One-Command Fix
Stop OpenClaw from spamming raw exec commands into Telegram chat by setting streaming.preview.toolProgress and streaming.progress.toolProgress to false. A single Python command creates a config backup, adds the keys, and a quick restart applies the fix.

Walking and Dictating with Claude Code: A Practical Setup
A developer shares their setup: walking 2-3 hours daily while dictating prompts to Claude Code via OpenAI Whisper, using a remote control or cowork chats.

KV Cache Quantization Issues in Local Coding Agents at High Context Lengths
A Reddit analysis identifies aggressive KV cache quantization as the cause of infinite correction loops and malformed JSON outputs in local coding agents like Qwen3-Coder and GLM 4.7 at 30k+ context lengths, recommending mixed precision or reduced context as workarounds.