OpenClaw Agent Broke at 65% Context: 683k Tokens, Zero Cache Reads on Ollama/GLM
A builder ran an OpenClaw agent named Francis in Discord for roughly 23 hours against GLM 5.3 Flash on Ollama Cloud, with a nominal 1,048,576 token context window. About 6 humans across several Discord channels talked to it while it handled tool calls, receipts, coding tasks, code reviews, and file reads. It held together through ~800 transcript events, then failed hard at roughly 65% of the configured context window.
The failure timeline
- ~683k tokens: Francis answered a message about a finished setup task with instructions from a completely unrelated job discussed ~19 hours earlier. Still coherent English, wrong moment.
- 2 minutes later: A long internal planning monologue about the stale task, then collapse into hundreds of repetitions of the word
design. - After that: checkmarks,
tool tool tool,toolResult, fake transcript, repeated numbers, and fragments of its own orchestration envelope. - ~688k tokens: final broken turns.
The session never recovered. Normal controls (/new, /reset, /stop) could not interrupt it, and the operator had to kill the OpenClaw session itself.
The part that actually matters: it reported success
No overflow error. No provider error. No timeout. From the harness's point of view, design design design was a perfectly valid model completion. Automatic compaction was expected around the point where roughly 25% of the window remained — it broke roughly 100k tokens before the safety net was due to fire.
Cache stats: cacheRead=0, cacheWrite=0
Every affected Ollama/GLM turn showed zero visible cache reads and zero cache writes. Cache telemetry was unavailable or zero, so OpenClaw had no evidence the repeated prefix was being reused. Around ten turns in the final 25 minutes each carried a prompt of roughly 680k tokens — about 6.1 million input tokens pushed through in that short window.
A separate agent session shortly before the main collapse recorded 6,912,253 input tokens, 28,956 output tokens, zero cache reads, zero compactions, no timeout, no provider error. Minutes later that agent claimed its conversation history contained fabricated people, tools, and channels, and admitted it could no longer tell which parts of its context were real.
The caching conversation
Minutes before the collapse, the operator asked Francis whether a caching layer was worth building, given all the repeated text moving through the system. Francis confidently said there was nothing useful to do because repeated context is already cheap at the provider via prompt/KV caching. That's a fine answer for OpenAI or Anthropic, where stable prefixes are cached and the telemetry proves it. It was false for the Ollama/GLM route in use — the agent was effectively telling the operator "this bit is cheap" while the harness re-sent ~680k tokens per turn.
The takeaway isn't "small model bad" — the agent did real work for nearly a full day first. The problem is the harness trusted a cache that wasn't there, had no compaction trigger before ~75%, and treated degenerate output as a successful completion. If you're running long-context agents against providers without verifiable cache telemetry, watch your input token accumulation, not just your context percentage.
📖 Read the full source: r/openclaw
👀 See Also

Non-developer builds word chain game in one day using Claude AI
A user with zero coding experience created a complete browser game in one session using Claude AI. The word chain game includes a 74k word dictionary, sound effects, design elements, and a mascot.

SeatBee.app Uses Claude AI for Wedding Seating Arrangements
SeatBee.app was built using Claude Code with Claude AI via OpenRouter to solve wedding seating chart problems. The AI handles constraint satisfaction for 150 guests with 20 rules, generates optimal seating in seconds, and understands social dynamics like creating buffer zones between people with messy breakups.

LLMs generate SQL queries to analyze terabytes of CI logs in seconds
Mendral's AI agent traced a flaky test to a dependency bump three weeks prior by writing its own SQL queries, scanning hundreds of millions of log lines across a dozen queries in seconds. The system handles 1.5 billion CI log lines weekly, compressed 35:1 in ClickHouse.

Multi-Agent AI Teams Using Context Baptism to Improve Code Reviews
A developer running 18 generations of AI agent teams discovered that agents who read letters and retrospectives from previous generations write significantly better code reviews than those who only read the code, calling this practice 'Context Baptism.'