Context Gateway: An Open-Source Proxy for Compressing AI Agent Context

What Context Gateway Does
Context Gateway is an agentic proxy that sits between AI coding agents (like Claude Code, OpenClaw, or Cursor) and the LLM API. When tool outputs like file reads or grep results dump thousands of tokens into the context window, the proxy compresses this content before it reaches the LLM. The motivation comes from research showing that long-context benchmarks experience steep accuracy drops as context grows—OpenAI's GPT-5.4 evaluation reportedly drops from 97.2% at 32k tokens to 36.6% at 1M tokens.
How the Compression Works
The system uses small language models (SLMs) that examine model internals and train classifiers to detect which parts of the context carry the most signal. When a tool returns output, compression happens conditioned on the intent of the tool call. For example, if an agent called grep looking for error handling patterns, the SLM keeps relevant matches and strips the rest. If the model later needs something that was removed, it can call expand() to fetch the original output.
Key Features and Setup
- Background compaction: Triggered at 85% window capacity, with summaries pre-computed so you don't wait for compaction
- Lazy-load tool descriptions: The model only sees tools relevant to the current step
- Spending caps: Control costs with budget limits
- Dashboard: Track running and past sessions
- Slack notifications: Get pinged when an agent is waiting on you
- Supported agents: Claude Code, Cursor, OpenClaw, or custom configurations
Getting Started
Install with:
curl -fsSL https://compresr.ai/api/install | sh
Then run context-gateway to launch an interactive TUI wizard that helps you:
- Choose an agent (claude_code, cursor, openclaw, or custom)
- Create/edit configuration including summarizer model and API key
- Enable Slack notifications if needed
- Set trigger threshold for compression (default: 75%)
The tool is open-source, built primarily in Go (90.9%), and maintained by Compresr, a YC-backed company. You can check compaction logs at logs/history_compaction.jsonl to see what's happening under the hood.
📖 Read the full source: HN LLM Tools
👀 See Also

Tracked 5 Biggest Claude Code SKILL.md Collections on GitHub — Sortable Table with Auto-Refresh
Built a sortable table of the top 5 skill-collection repos (totaling 125k stars) with star counts and skill counts, auto-refreshed by a /workflows:skill-collections command.

Local AI Agent Achieves Sub-Second STT and TTS Latency with Open-Source Servers
A developer achieved ~0.2s STT latency using Whisper large-v3-turbo with hybrid thread-managed GPU architecture and ~250ms TTS latency with Coqui-TTS optimized for low-latency synthesis. Both implementations are fully self-hosted and open-sourced.

Browser CLI: A Token-Efficient Browser Automation Tool for AI Coding Agents
Browser CLI is a persistent headless Chromium daemon that provides browser automation via plain Bash commands, achieving ~95% token savings compared to Playwright MCP by reducing calls from ~1,500 tokens to ~75 tokens.

Crispy VS Code Extension Adds Agent Memory and Multi-Agent Features for Claude and Codex
Crispy is an open-source VS Code extension that wraps Claude Code and Codex CLIs with a GUI, adding local agent memory with semantic search, multi-agent sessions, conversation forking, and dedicated tool views. It runs on Linux, macOS, and Windows under MIT license.