Cognithor: A Local-First Agent OS with PGE Trinity Architecture

Cognithor is a fully local, autonomous Agent OS developed over one year through 16 distinct phases. The project emphasizes deliberate architecture, documented decisions, and substantial test coverage, distinguishing it from what the creator calls "vibe-coded" AI projects.
Core Architecture: PGE Trinity
Every task in Cognithor flows through a three-gate system: Planner → Gatekeeper → Executor. The Gatekeeper is deterministic, enforcing policy before execution rather than after, creating a control layer beyond simple agent chaining.
Technical Specifications
- Codebase: >118,000 LOC source, >108,000 LOC test
- Testing: 11,609+ tests with 89% coverage, 0 lint errors
- LLM Support: 16 providers including Ollama, LM Studio, Anthropic, OpenAI, Gemini, and 11 others
- Channels: 17 interfaces including Telegram, Discord, Slack, WhatsApp, Signal, Voice, CLI, and WebUI
- Tools: 123 MCP tools
- Features: Computer Use, Deep Research v2 (25-round iterative), SSH remote execution, VS Code extension
- Memory: 5-tier cognitive memory system
- Security: GDPR-compliant with Ed25519-signed audit trail
Local-First Implementation
The system operates with no cloud requirements and no mandatory API keys. All data remains on the user's machine, with Ollama or LM Studio running the brain. Cloud providers are available as opt-in alternatives.
Development Phases
The 16 completed phases include foundation (PGE, MCP, CLI), multi-agent collaboration, GDPR toolkit, distributed workers, and a Flutter Command Center. Each phase is documented, tested, and shipped.
The project is developed primarily by one developer with assistance from a tester in Budapest who validates the system on fresh machines. The developer notes that "AI writes the code. I engineer the system."
The GitHub repository is available at Alex8791-cyber/cognithor, with v1.00.0 expected to be released soon.
📖 Read the full source: r/LocalLLaMA
👀 See Also

OmniCoder-9B fine-tune shows strong performance for agentic coding on 8GB VRAM systems
A Reddit user tested OmniCoder-9B, a fine-tune of Qwen3.5-9B on Opus traces, with OpenCode and reported 40+ tokens per second speeds using Q4_K_M GGUF quantization at 100k context length on an 8GB VRAM system.

Dual-model architecture reduces token consumption by half for long conversations
A developer built a dual-model system where a small 'subconscious' model compresses conversation history in the background, allowing the main model to work with a curated ~35K context instead of 120K tokens of raw history. This architecture cuts token consumption roughly in half for sustained project work.

context-link v1.0.0: Local MCP server reduces Claude Code token usage by 91%
context-link v1.0.0 is a local MCP server that indexes codebases with Tree-sitter to serve Claude only the exact symbols, dependencies and structure needed, reducing token usage by 91% in specific cases and 70-80% across full tasks.

Maestro v1.5.0 adds Claude Code support for multi-agent orchestration
Maestro v1.5.0, an open-source multi-agent orchestration platform, now runs as a native plugin on Claude Code in addition to Gemini CLI. The update includes deeper design planning, a 42-step orchestration backbone, agent capability enforcement, and security hardening.