DeepSeek V4 pricing reality check: 178x cheaper cached tokens vs Opus, but capability lag acknowledged

DeepSeek V4 launched with pricing so low a Reddit user checked the math. Here are the verified numbers:
Pricing breakdown
- V4-Pro standard input: $0.145 per million tokens. Opus 4.7 input: ~$5 per million. Ratio: 34x.
- With 75% promotional discount (through end of May): V4-Pro input drops to $0.036 per million — 138x cheaper than Opus.
- Cache hit pricing: V4-Pro is $0.0036 per million. Opus cached is $0.625 per million. Ratio: 173x.
The catch
As the original post notes, DeepSeek admits V4 is three to six months behind GPT-5.4 and Gemini 3.1 Pro on capability. You're not getting frontier quality at frontier-divided-by-178 — you're getting last summer's frontier quality.
What this means for agentic workflows
For agentic loops with heavy caching (system prompts, tool definitions), the cache hit discount is the real story. Reusable system prompts become essentially free. The key unknown: whether the claimed 1M context window holds up under real workloads or degrades to a usable 200K, as seen with many large-window models.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Practical Enhancements in Claude Opus 4.6: Memory Upgrade
Claude Opus 4.6 features a significant upgrade with a 1 million token context, enhancing memory retention and performance in complex tasks.

Claude Code System Prompt Assembly and Structure Revealed
A source map leak in Claude Code's npm package exposed the system prompt assembly flow, showing static prefix sections followed by dynamic session-specific content, with three identity variants and detailed execution guidelines.

Diagnosing Operational Drift and Task Amnesia in OpenClaw with Gemini 2.5 Flash on Proxmox
OpenClaw users report issues with persistent workflows on a Proxmox VM, citing operational drift and task amnesia. Despite stable performance in one-off tasks, the Gemini 2.5 Flash model struggles with automation and memory in this setup.

AI Tools Increase Engineering Workload and Shift Professional Roles
A February 2026 Harvard Business Review study found 83% of workers reported increased workload from AI tools, with 62% experiencing burnout. The article describes how AI has shifted engineering roles from writing code to reviewing AI-generated code.