ProofShot: CLI for AI Agents to Verify UI Code with Browser Recording

What ProofShot Does
ProofShot is a CLI tool that gives AI coding agents visual verification capabilities. It allows agents to see what the UI they build actually looks like in the browser, detect layout issues, and capture console errors.
How It Works
The tool operates through three main commands:
proofshot start --run "npm run dev" --port 3000- Launches your dev server, opens headless Chromium, and starts recording video- Your AI agent then executes actions like
proofshot exec navigate "http://localhost:3000"andproofshot exec screenshot "homepage"to navigate, click, fill forms, and take screenshots proofshot stop- Collects errors, stops recording, trims dead time, and generates proof artifacts
Output and Features
ProofShot generates a standalone HTML file containing:
- Video playback of the browser session synced with an action timeline
- Screenshots taken during the session
- Element labels for each action
- Browser console errors captured during the session
- Server logs scanned with pattern matching for JavaScript, Python, Go, Rust, and other languages
- PR-ready artifacts including SUMMARY.md and formatted output for pull requests
- Visual diff comparison against baselines
Technical Details
The tool is:
- Built on agent-browser from Vercel Labs (described as "far better and faster than Playwright MCP")
- Not a testing framework - the agent doesn't decide pass/fail, it just provides evidence
- Agent-agnostic - works with Claude Code, Cursor, Codex, Gemini CLI, Windsurf, and any MCP-compatible agent
- Packaged as a skill so AI agents know exactly how it works
- Open source with MIT license
Installation and Setup
$ npm install -g proofshot
$ proofshot install
The tool automatically trims dead time from recordings, so you see only what the agent actually did, not idle waiting periods.
📖 Read the full source: HN LLM Tools
👀 See Also

Anchormd: A Tool for Managing Context Across Claude AI Sessions
Anchormd is an open-source tool that addresses context loss in Claude AI sessions by indexing curated markdown plans into a searchable knowledge graph. It allows agents to load project overviews at session start and query for specific details as needed.

Driftwatch V3 Released: AI-Assisted Codebase Monitoring Tool
Driftwatch V3 is now available as a public repository after a 5-6 day build involving approximately 9,000 lines of code and $160 in API credits. The in-browser tool tracks markdown file issues, flags contradictory instructions, and provides cost tracking with recommendations.

Local LLM Performance Benchmarks on Mac Mini with OpenClaw and LM Studio
A Reddit user posted performance figures for running the Unsloth gpt-oss-20b-Q4_K_S.gguf model locally on a Mac Mini with 32GB RAM, achieving 34 tokens/second with a 0.7 second time to first token using OpenClaw 2026.3.8 and LM Studio 0.4.6+1.

Open Source SQLite-Based Persistent Memory System for Claude
A developer has released memchat, a GPL-licensed local system that extracts knowledge from Claude sessions at checkpoints, stores it in SQLite, and reassembles it for new sessions to maintain context across conversations.