ProofShot: CLI for AI Agents to Verify UI Code with Browser Recording

What ProofShot Does
ProofShot is a CLI tool that gives AI coding agents visual verification capabilities. It allows agents to see what the UI they build actually looks like in the browser, detect layout issues, and capture console errors.
How It Works
The tool operates through three main commands:
proofshot start --run "npm run dev" --port 3000- Launches your dev server, opens headless Chromium, and starts recording video- Your AI agent then executes actions like
proofshot exec navigate "http://localhost:3000"andproofshot exec screenshot "homepage"to navigate, click, fill forms, and take screenshots proofshot stop- Collects errors, stops recording, trims dead time, and generates proof artifacts
Output and Features
ProofShot generates a standalone HTML file containing:
- Video playback of the browser session synced with an action timeline
- Screenshots taken during the session
- Element labels for each action
- Browser console errors captured during the session
- Server logs scanned with pattern matching for JavaScript, Python, Go, Rust, and other languages
- PR-ready artifacts including SUMMARY.md and formatted output for pull requests
- Visual diff comparison against baselines
Technical Details
The tool is:
- Built on agent-browser from Vercel Labs (described as "far better and faster than Playwright MCP")
- Not a testing framework - the agent doesn't decide pass/fail, it just provides evidence
- Agent-agnostic - works with Claude Code, Cursor, Codex, Gemini CLI, Windsurf, and any MCP-compatible agent
- Packaged as a skill so AI agents know exactly how it works
- Open source with MIT license
Installation and Setup
$ npm install -g proofshot
$ proofshot install
The tool automatically trims dead time from recordings, so you see only what the agent actually did, not idle waiting periods.
📖 Read the full source: HN LLM Tools
👀 See Also

Building an Autonomous Research Agent with C# and Local LLMs
A C# research agent automates URL processing with local LLMs using Ollama and llama3.1:8b, generating structured markdown reports from web searches.

MemAware Benchmark Tests AI Memory Beyond Keyword Search
MemAware is a benchmark with 900 questions across 3 difficulty levels that tests whether AI assistants with memory can surface relevant context when queries don't hint at it. Results show BM25 search scored 2.8% vs 0.8% with no memory, while vector search drops to 0.7% on cross-domain connections.

Meta Ads MCP OAuth Works But Most Ad Accounts Not Enabled Yet
Meta Ads MCP OAuth flow works and loads 29 tools, but ads_get_ad_accounts returns is_ads_mcp_enabled: false with a message that the feature is gradually rolling out.

MCP Server Tracks Known Bugs in Dev Tools to Improve LLM Recommendations
nanmesh-mcp is an MCP server that crawls GitHub Issues, Stack Overflow, and Reddit to track real problems in 57 development tools, providing LLMs with current bug data before making library recommendations.