ProofShot CLI Gives AI Coding Agents Browser Verification Capabilities

ProofShot: Browser Verification for AI Coding Agents
ProofShot is an open-source, agent-agnostic CLI that gives AI coding agents the ability to verify UI features they build by recording browser sessions, capturing screenshots, and collecting errors. It addresses the problem where agents write code but can't see what it actually looks like in the browser or detect layout issues and console errors.
How It Works
The tool follows a three-step workflow: start, test, stop. The AI agent drives the browser using agent-browser commands while ProofShot records the session.
Basic usage:
proofshot start --run "npm run dev" --port 3000
# agent navigates, clicks, takes screenshots
proofshot stop
Detailed workflow example:
# 1. Start — open browser, begin recording, capture server logs
proofshot start --run "npm run dev" --port 3000 --description "Login form verification"
2. Test — the AI agent drives the browser
agent-browser snapshot -i # See interactive elements
agent-browser open http://localhost:3000/login # Navigate
agent-browser fill @e2 "[email protected]" # Fill form
agent-browser click @e5 # Click submit
agent-browser screenshot ./proofshot-artifacts/step-login.png # Capture proof
3. Stop — bundle video + screenshots + errors into proof artifacts
proofshot stop
Key Features
- Works with any AI coding agent that can run shell commands (Claude Code, Cursor, Codex, Gemini CLI, Windsurf, GitHub Copilot, etc.)
- Packaged as a skill so AI agents understand how to use it
- Built on agent-browser from Vercel Labs (described as "far better and faster than Playwright MCP")
- Not a testing framework — doesn't decide pass/fail, just provides evidence
- Generates self-contained HTML files with video, screenshots, and logs
- Can upload artifacts to GitHub PRs as inline comments with
proofshot pr
Installation and Setup
npm install -g proofshot
proofshot install
The first command installs the CLI and agent-browser (with headless Chromium). The second detects your AI coding tools and installs the ProofShot skill at user level — works across all projects automatically.
Output Artifacts
Each session produces a timestamped folder in ./proofshot-artifacts/ containing:
session.webm— Video recording of the entire sessionviewer.html— Standalone interactive viewer with scrub bar, timeline, and Console/Server log tabsSUMMARY.md— Markdown report with errors, screenshots, and videostep-*.png— Screenshots captured at key momentssession-log.json— Action timeline with timestamps and element dataserver.log— Dev server stdout/stderr (when using--run)console-output.log— Browser console output
Available Commands
proofshot install— Detect AI coding tools and install ProofShot skillproofshot start— Start verification session with browser, recording, error captureproofshot stop— Stop recording, collect errors, generate proof artifactsproofshot exec— Pass-through command
The tool is completely free and open source, with no vendor lock-in or cloud dependency. It's designed for developers who use AI agents to build UI features and want to verify the results without manually opening the browser each time.
📖 Read the full source: HN AI Agents
👀 See Also

Clawback: Hooks-based implementation of leaked Claude verification loops
Clawback is a GitHub project that reimplements the verification loops from the Claude source map leak as mechanical hooks instead of prompts. It includes stop hooks, PreToolUse, PostToolUse, and PostCompact hooks that can't be skipped by the model under context pressure.

Google PM Open-Sources Always On Memory Agent with SQLite Storage, No Vector DB
Google senior AI product manager Shubham Saboo has open-sourced an Always On Memory Agent that stores structured memories in SQLite instead of using vector databases, running on Gemini 3.1 Flash-Lite with scheduled memory consolidation every 30 minutes.

Phalanx CLI coordinates multiple AI agents for automated code-review cycles
A developer built Phalanx, a CLI tool that coordinates AI agents from different providers: Codex handles coding, Claude Opus performs code review, and Claude Sonnet orchestrates the loop. A companion tool called Codebones compresses repositories to structural maps to reduce token usage.
Forking OpenClaw with a Custom LLM: A Local-Only Setup Guide
A Reddit user forked OpenClaw to run entirely on local models, creating a fully customized 'JARVIS' with zero cloud dependency. The process took minutes, not hours, and runs two versions side-by-side on an M2 Ultra.