AMD Threadripper Halo Station: 96-Core AI Workstation with MI350P
At IFA 2026, AMD announced the Threadripper Halo Station, a workstation packing a 96-core Threadripper Pro 9995WX and dual liquid-cooled MI350P accelerators, designed to run trillion-parameter models. AMD claims it's "the most powerful workstation in the world," though no OEM partners have been announced yet.
Core specs
- CPU: Threadripper Pro 9995WX — 96 Zen 5 cores, 192 threads, boost up to 5.4 GHz, 384 MB L3 cache, 350W TDP.
- Accelerators: Dual liquid-cooled Instinct MI350P, with a path to four.
- Memory: 2 TB DDR5 system RAM.
- HBM: 288 GB HBM3E on the GPUs, with up to 576 GB supported (presumably with four accelerators).
The machine essentially reconfigures a server tray into a tower, swapping an EPYC host for the Threadripper. AMD has not detailed the full system design or pricing, but street price for just the CPU, GPUs, and 2 TB of RAM is over $100,000. A fully configured unit—with storage, power, and cooling—could easily exceed $150,000.
AMD hasn't announced which OEMs will build and ship the system. Given the price and target workload (running trillion-parameter models locally), this is squarely aimed at AI research labs and enterprise teams that need on-prem inference without cloud round-trips.
📖 Read the full source: HN AI Agents
👀 See Also

SAP Freezes Most Travel and Hiring Due to AI's Soaring Cost — Except AI Itself
SAP suspended most travel and hiring due to AI's soaring cost, with exceptions only for AI-related hires and travel, according to an internal email obtained by 404 Media.

New AI Tutor Achieves 0.71-1.30 SD Effect Size in Dartmouth Course
A new AI tutor for a Dartmouth introductory CS course showed learning gains of 0.71 to 1.30 standard deviations compared to a control group.

GLM-5.1 Released with Coding Performance Matching Claude Opus 4.5
Zhipu AI's GLM-5.1 model is now available to all Coding Plan users, achieving 77.8 points on SWE-bench-Verified and 56.2 points on Terminal Bench 2.0. The model features a 200K context window, 128K max output, and 744B parameters with 40B activated.

Benchmarks Show Distilled Models Match Frontier LLMs on Structured Tasks at 10x Lower Cost
A comprehensive comparison of small distilled Qwen3 models (0.6B to 8B) against frontier LLMs shows distilled models match or beat mid-tier frontier models on 6 out of 9 tasks at dramatically lower cost, with Text2SQL achieving 98.0% accuracy at $3/M requests versus $378 for Claude Haiku.