SigMap v8.9: Deterministic Context Layer Cuts Token Use 97% for AI Coding Agents

SigMap v8.9 is a deterministic, verifiable grounding layer for AI code work. It claims a 97.0% token reduction, 87.8% hit@5 retrieval accuracy (vs 13.6% random), and 49.2% fewer prompts per task (1.44 vs 2.84). The tool is zero-dependency, fully offline, and works with TypeScript, Python, Go, Rust, Java, Kotlin, Ruby, PHP, Swift, C#, C++, Dart, Scala, Vue, Svelte, GraphQL, SQL, Terraform, R, GDScript, and more.
Core workflow: ask, validate, judge, learn
The workflow goes beyond simple context shrinking:
- Generate a compact signature map once with
npx sigmap. - Ask for files specific to the current task:
sigmap ask "explain the auth flow". This outputs a ranked file list and.context/query-context.mdready to paste. - Validate coverage:
sigmap validate --query "auth login token"checks if context is sufficient. - Judge grounding:
sigmap judge --response response.txt --context .context/query-context.mdscores whether the answer is grounded in the code.
MCP and IDE integration
SigMap is MCP-ready and works with Copilot, Claude Code, Cursor, Windsurf, Codex, OpenCode, and Gemini CLI. The v8.9.1 release adds a squeeze_output MCP tool and squeeze --response CLI flag to compress noisy stack traces, CI logs, or JSON payloads deterministically mid-session — the 19th MCP tool.
Benchmark results
Latest saved benchmark (v8.9.1, July 2026):
| Metric | Without SigMap | With SigMap |
|---|---|---|
| Task success proxy | 10% | 67.8% |
| Prompts per task | 2.84 | 1.44 |
| Retrieval hit@5 | 13.6% | 88% |
| Overall token reduction | — | 97.0% |
| GPT-4o overflow repos | 16/21 | 0/21 |
Performance spans 21 repos and 90 real coding tasks.
Quick start
npx sigmap
sigmap ask "explain the auth flow"
# Outputs ranked file list + .context/query-context.md
# Paste context into AI assistant
For teams or CI, SigMap offers configurable strategies and generalization for monorepos. It also has a dedicated guide for open-source agents (OpenCode, Aider, Cline) and local LLMs (Ollama, llama.cpp, vLLM) — zero cost, full privacy.
📖 Read the full source: HN LLM Tools
👀 See Also

GLM-5-Turbo Shows Low Tool Call Error Rate in User Testing
The z-ai/glm-5-turbo model demonstrates a 0.57% average tool call error rate in testing, significantly lower than GLM-5's ~3% rate. A user reported successfully using it with a CLI tool to write a 97,000-word fantasy novel with minimal issues.

Open-source tool enables Claude to control Unreal Engine directly
soft-ue-cli is a Python tool with a C++ plugin that allows Claude Code and Claude Desktop to execute commands in Unreal Engine without editor interaction, featuring 60+ operations including blueprint editing, actor spawning, and performance profiling.

FUTO Swipe: Open-Source Swipe Typing Models Match Big Tech Accuracy
FUTO releases open-source swipe typing models and a 1M swipe dataset. Encoder (635K params) + ContextLM (1.5M) + decoder (304K) achieve ~4% top-4 fail rate. Fully offline in FUTO Keyboard.

Debugging Claude Code's Build-Check Logic: Why Name Search Fails and Structural Footprint Search Fixes It
Claude Code told a user 'feature not built' four times in one session — all wrong. The fix: replace name-based search with structural footprint search (routes, schemas, registered tools). Practical rule shared.