SubQ: First Fully Subquadratic LLM with 12M-Token Context and 95% RULER Accuracy

Subquadratic has released SubQ 1M-Preview, the first fully subquadratic large language model, where compute scales linearly with context length — not quadratically as with transformers. This eliminates the need for RAG systems and chunking workarounds for long-context tasks. The research model supports up to 12 million tokens, with a 1M-token production model available in early access.
Key Features
- Subquadratic attention: Reduces attention compute by ~1,000x compared to frontier transformer models at 12M-token context, per the source.
- SubQ Code: CLI-based coding agent that loads entire codebases into a single context window. No multi-agent orchestration needed — plans, executes, and reviews across a full repository in one pass.
- SubQ Search: Long-context search tool offering Deep Research capabilities at chatbot speed.
- API: Full-context API for developers and enterprise teams.
Benchmarks
All results were verified by a third party (source does not specify the firm):
- RULER 128K: 95% accuracy — compared to Claude Opus 4.6 at 94.8%.
- MRCR v2 (multi-piece retrieval & reasoning): Production model scores 65.9; research model scores 83. Reference: Claude Opus 4.7 = 32.2, GPT 5.5 = 74, Gemini 3.1 Pro = 26.3.
- SWE-Bench Verified: 81.8% — compared to Opus 4.6 (80.8) and Deepseek 4.0 Pro (80.0).
- Attention speed: SubQ Sparse Attention is 52× faster than FlashAttention in architecture-level comparison, using 63% less compute.
Architecture Details
The model uses a fundamentally redesigned attention mechanism built from first principles to be subquadratic. It leverages linear attention, state space model ideas, and sparse attention — but unlike prior attempts, maintains frontier-level accuracy. The team includes PhDs from Meta, Google, Oxford, BYU, ByteDance, Adobe, and Cambridge.
Availability
Private beta starts today (May 5, 2026). Access to API, SubQ Code CLI, and SubQ Search. SWE-Bench score indicates strong coding performance for AI coding agents like OpenClawRadar's readers.
📖 Read the full source: HN AI Agents
👀 See Also

Claude Tops App Store Charts Amid Government Standoff
Anthropic's Claude app jumped from 42nd to 1st place on the US App Store's Top Downloaded charts, with ChatGPT and Gemini taking second and third. The surge follows a public disagreement between Anthropic and the US government over military and surveillance use of AI technology.
There's No Limit to How Bad Code Can Get
A former Amazon engineer shares why the 'sinking ship' metaphor for bad codebases is misleading: business ships sink, but software never hits bottom. Code can always get worse.

Pentagon to adopt Palantir AI as core US military system
The Pentagon plans to adopt Palantir's AI technology as a core system for the US military, according to a memo. The Reuters article generated 47 points and 2 comments on Hacker News.

Talkie: A 13B LLM Trained Exclusively on Pre-1931 Text, Using Claude as a Judge in RL Training
Researchers released Talkie, a 13B LLM trained only on text published before 1931 (no internet, no WWII data). Claude Sonnet 4.6 was used as the judge in its online DPO reinforcement learning pipeline, and Claude Opus 4.4 generated synthetic multi-turn conversations for fine-tuning. The model can write Python code from a few in-context examples despite zero modern code in training.