Strands Decider 2B: A 2B-Parameter Decision Model for Agentic Workflows
Strands Labs (AWS) released Strands Decider 2B, a 2-billion-parameter open-source decision model optimized for fast local experimentation and agentic workflows. Unlike general-purpose LLMs, decision models are designed to pick between a fixed set of options (e.g., "Is 'turn on the lights' about the coffee machine? Yes or no.") or assign simple numerical scores (e.g., sentiment between 0 and 1).
What it is and why it matters
Decision models trade flexibility for speed and reliability. They always return an answer from the provided options, run with very low latency, and emit a calibration score ("how sure can I be this is correct?") that frontier LLM APIs don't expose. The tradeoff: they're worse at complex reasoning and can't generate text, so they're unsuited for coding, chatbots, or summarization.
Strands positions this as a new class of "system one" models, following TypeSafe AI's launch of Jev earlier this month. The intent is to drive agentic decision steps in the Strands Harness SDK.
Architecture
The model takes a pre-trained Qwen3.5-2B torso, strips the LM head (removing text generation), and replaces it with a pointer head that scores each offered option. The head compares the hidden state at each option position against the hidden state at the <answer> position. The head is small — just over 1M parameters — and the torso is fine-tuned with a rank-16 LoRA adapter.
This is v19 of the architecture. The first iteration used a slot head, which performed significantly worse; all iteration notes are in the repo.
Benchmarks
Strands measured accuracy on JevBench's public set and calibration via Brier score on the same set:
- Accuracy: 3rd of 33 in the 2B class
- Accuracy excluding just-over-2B models: 1st of 30
- Latency (median): ~115ms on an Nvidia RTX 3090, ~153ms for small tasks on an M3 MacBook
- Latency scales roughly linearly with task size
Availability
- Code on GitHub:
strands-decider-2b - Weights on Hugging Face
- Includes all training data and scripts
- Runs on local CPU or GPU — no API calls required
Who it's for
Developers building agentic workflows who need fast, calibrated yes/no or classification decisions at the edge, and researchers who want a small, hackable base for decision-model experiments.
📖 Read the full source: HN AI Agents
👀 See Also

OpenClaw .NET: NativeAOT Port with JSON-RPC Bridge for Existing Plugins
OpenClaw .NET is a C# port of OpenClaw that compiles to a ~23MB NativeAOT binary, eliminating JIT warmup and Node runtime overhead while maintaining compatibility with existing TypeScript/JavaScript plugins through a built-in JSON-RPC bridge.

JetBrains Introduces Plugin for Modern Go Code with AI Agents Junie and Claude Code
JetBrains has released a plugin for AI agents Junie and Claude Code, enhancing their ability to generate modern Go code by adhering to the latest Go features and best practices.

Depct tool collects runtime data to help Claude debug production issues
Depct is a tool that collects runtime instrumentation from Node.js apps, builds graphs from the data, and feeds it to Claude via AWS Bedrock to help debug intermittent production failures. It also generates architecture diagrams and dependency maps from runtime behavior.

Harnessing Claude Code for Bot Consultancy: A Deep Dive
Exploring the integration of Claude Code within bot development to enhance functionality through AI consultancy, as shared by an enthusiast on r/clawdbot.