3 weeks of OpenClaw: token costs, loops, and compaction — lessons from the trenches

A developer on r/openclaw shared hard-won lessons after three weeks with OpenClaw. The post covers five major pain points and their fixes — practical advice for anyone stuck in the agent-configuration grind.
Don't use Opus for everything
The biggest money waste: running Opus on trivial tasks like heartbeat checks and cron pings. The user switched to glm-5.1 for routine work and only uses sonnet 4.6 for tasks that require reasoning. This cut token costs by roughly two-thirds.
Agents loop and forget by default
Out-of-the-box agents loop, forget decisions, and ask bizarre questions. The fix: write custom rules including anti-loop instructions, context summaries, and verification steps that make the agent confirm its actions before asking for more input. The user emphasizes this is tedious but essential — it's what separates a working agent from a broken one.
Start small, add features one at a time
Trying to wire up email, WhatsApp, web scraping, and cron simultaneously broke everything. The user backed up, started with email summaries only, got that solid, then added each feature incrementally. Obvious advice, but easy to ignore when excited.
Compaction destroys long-term context
OpenClaw's context compaction gradually erases decisions made days ago. Workaround: dump important info into workspace docs, maintain decision logs, and feed the agent reference material before each session. It's annoying but makes a night-and-day difference in agent memory.
Consider Autoclaw for setup if not technical
For users overwhelmed by initial configuration, Autoclaw offers a one-click installer with preloaded skills. The user found it helpful to avoid fighting with installation issues.
The user's final warning: those “my agent built a full app overnight” posts come from people who spent weeks tuning their config first. Don't compare your day three to their month three.
📖 Read the full source: r/openclaw
👀 See Also

Claude Prompt Codes Retested: L99 Sharper, OODA Narrower, ARTIFACTS Faded, and 3 New Codes to Use
A 6-month retest of L99, OODA, and ARTIFACTS prompt codes on Claude shows L99 sharper on Sonnet 4.6/Opus 4.7, OODA failing on strategic prompts, ARTIFACTS unnecessary for code, and three new codes (/skeptic, /blindspots, /decompose) earning daily use. Stack no more than 2 codes.

Auth 400 Error Fix: Using Python's mnemonic Package to Avoid BIP39 Filter Triggers
A Reddit user identified that Anthropic's content filter triggers a 400 error when AI agents attempt to write the full BIP39 wordlist (2048 standardized English words) into Python code. The solution is to use the mnemonic Python package instead, which contains the wordlist internally.

Verification Harness Fixes Claude's Plan Execution Problem
A developer built a 30-50 line bash or Python verification layer that checks whether Claude actually executes each step of its own plans by verifying artifacts like file existence, API responses, and config changes.

Reddit user shares common Claude Code prompting mistakes with fixes
A developer using Claude for Node.js backend work identified 10 common prompting mistakes after months of use, including missing validation requirements and treating Claude as one-shot tool. They created a visual guide with fixes for each issue.