OpenClaw Agent Broke at 65% Context: 683k Tokens, Zero Cache Reads on Ollama/GLM

✍️ OpenClawRadar📅 Published: September 30, 2026🔗 Source
Ad

A builder ran an OpenClaw agent named Francis in Discord for roughly 23 hours against GLM 5.3 Flash on Ollama Cloud, with a nominal 1,048,576 token context window. About 6 humans across several Discord channels talked to it while it handled tool calls, receipts, coding tasks, code reviews, and file reads. It held together through ~800 transcript events, then failed hard at roughly 65% of the configured context window.

The failure timeline

  • ~683k tokens: Francis answered a message about a finished setup task with instructions from a completely unrelated job discussed ~19 hours earlier. Still coherent English, wrong moment.
  • 2 minutes later: A long internal planning monologue about the stale task, then collapse into hundreds of repetitions of the word design.
  • After that: checkmarks, tool tool tool, toolResult, fake transcript, repeated numbers, and fragments of its own orchestration envelope.
  • ~688k tokens: final broken turns.

The session never recovered. Normal controls (/new, /reset, /stop) could not interrupt it, and the operator had to kill the OpenClaw session itself.

The part that actually matters: it reported success

No overflow error. No provider error. No timeout. From the harness's point of view, design design design was a perfectly valid model completion. Automatic compaction was expected around the point where roughly 25% of the window remained — it broke roughly 100k tokens before the safety net was due to fire.

Ad

Cache stats: cacheRead=0, cacheWrite=0

Every affected Ollama/GLM turn showed zero visible cache reads and zero cache writes. Cache telemetry was unavailable or zero, so OpenClaw had no evidence the repeated prefix was being reused. Around ten turns in the final 25 minutes each carried a prompt of roughly 680k tokens — about 6.1 million input tokens pushed through in that short window.

A separate agent session shortly before the main collapse recorded 6,912,253 input tokens, 28,956 output tokens, zero cache reads, zero compactions, no timeout, no provider error. Minutes later that agent claimed its conversation history contained fabricated people, tools, and channels, and admitted it could no longer tell which parts of its context were real.

The caching conversation

Minutes before the collapse, the operator asked Francis whether a caching layer was worth building, given all the repeated text moving through the system. Francis confidently said there was nothing useful to do because repeated context is already cheap at the provider via prompt/KV caching. That's a fine answer for OpenAI or Anthropic, where stable prefixes are cached and the telemetry proves it. It was false for the Ollama/GLM route in use — the agent was effectively telling the operator "this bit is cheap" while the harness re-sent ~680k tokens per turn.

The takeaway isn't "small model bad" — the agent did real work for nearly a full day first. The problem is the harness trusted a cache that wasn't there, had no compaction trigger before ~75%, and treated degenerate output as a successful completion. If you're running long-context agents against providers without verifiable cache telemetry, watch your input token accumulation, not just your context percentage.

📖 Read the full source: r/openclaw

Ad

👀 See Also