Cerebras releases Step-3.5-Flash-REAP models with 40% memory reduction

What this is
Cerebras has released Step-3.5-Flash-REAP models, which are memory-efficient compressed variants of their larger models. These are smaller versions designed for what the source calls "potato setups," though the 121B parameter model still requires significant resources.
Key details from the source
The models are available on Hugging Face:
The Step-3.5-Flash-REAP-121B-A11B model is compressed from 196B to 121B parameters, representing a 40% memory reduction while maintaining near-identical performance to the full model.
The compression uses REAP (Router-weighted Expert Activation Pruning), described as "a novel expert pruning method that selectively removes redundant experts while preserving the router's independent control over remaining experts."
Features and capabilities
- Near-lossless performance: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 196B model
- 40% memory reduction: Compressed from 196B to 121B parameters, lowering deployment costs and memory requirements
- Preserved capabilities: Retains all core functionalities including code generation, math & reasoning, and tool calling
- Drop-in compatibility: Works with vanilla vLLM - no source modifications or custom patches required
- Optimized for real-world use: Particularly effective for resource-constrained environments, local deployments, and academic research
The source notes that while these are "smaller versions," the 121B model still requires a fairly powerful setup despite the compression.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Revolutionize API Monitoring Across Providers with onWatch
Discover how onWatch, a powerful new tool, streamlines tracking your AI API quota usage across multiple providers, ensuring you stay within limits and optimize resource allocation.
How Claude's Text Watermarking Works
Claude's upcoming watermarking uses a key and preceding words to alter random choices without affecting output quality. It's designed to comply with the EU AI Act.

Anthropic Removes Claude Code from Pro Subscription for New Users in Test
Anthropic temporarily removed access to Claude Code from its $20/month Pro subscription plan for new users, changing website pricing pages and support documents before reversing the changes. The company described it as a 'small test of 2% of new prosumer signups.'

Claude Code v2.1.150 Adds Remote System Prompt Injection via Network
Claude Code v2.1.150 fetches system prompts from Anthropic servers at startup and every 60 seconds via a GrowthBook feature flag, allowing remote injection—bypassed with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1.