Claude Code developer acknowledges adaptive thinking flaw, provides workaround

Boris Charny, the creator of Claude Code at Anthropic, publicly engaged with developers on Hacker News about performance issues reported since February. Initially attributing problems to user settings, he shifted his position after examining bug transcripts.
Initial position: Settings issue
Charny's first explanation pointed to two intentional changes: hiding the thinking process (a UI change) and lowering the default effort level. The implicit message was that performance hadn't degraded - users were just experiencing the new, lower-cost default. He suggested changing settings back to /effort high for previous performance levels.
Shift to acknowledgment
When confronted with evidence from users already using the highest effort settings and still experiencing problems, Charny analyzed bug reports and moved from general settings explanations to specific technical diagnosis.
Final position: Specific flaw identified
Charny explicitly validated users' experiences, conceding that the "adaptive thinking" feature is "under-allocating reasoning." He confirmed this wasn't related to effort defaults, as telemetry showed affected sessions were sending effort=high on every request.
His final message stated: "The data points at adaptive thinking under-allocating reasoning on certain turns — the specific turns where it fabricated (stripe API version, git SHA suffix, apt package list) had zero reasoning emitted, while the turns with deep reasoning were correct."
Workaround and investigation
Charny provided an interim workaround: CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1 forces a fixed reasoning budget instead of letting the model decide per-turn. He noted the model team is investigating the issue.
The discussion demonstrates direct engagement between Anthropic technical staff and external developers, with transparency about technical issues affecting Claude Code's performance.
📖 Read the full source: r/ClaudeAI
👀 See Also

Reddit user reports 18.8 tok/s CPU inference with Qwen 3 30B Q4 on Zen 4
A user on r/LocalLLaMA tested Qwen 3 30B Q4 on CPU and achieved 18.8 tokens per second with a Zen 4 processor and DDR5 memory, significantly exceeding expectations of 3-5 tok/s.

llama.cpp Q8_0 quantization gets 3.1x speedup on Intel Arc GPUs with SYCL reorder fix
A fix to llama.cpp's SYCL backend brings Q8_0 quantization on Intel Arc GPUs from 21% to 66% of theoretical memory bandwidth, achieving 15.24 tokens/second versus 4.88 tokens/second previously on an Arc Pro B70 with Qwen3.5-27B.

Seven Ways to Avoid Losing Your Job to AI – Tyler Cowen's Practical Guide
Tyler Cowen outlines seven principles, including seeking messy jobs and being wary of remote work, to protect your career against AI competition.
OpenClaw Updates Break Setups: What Users Are Doing About It
A user managing 5 sites with OpenClaw, Claude Code, and Codex shares that updates often break their working setup, prompting a cautious approach to upgrading.