Scalpel v2.0: Codebase Scanner and AI Agent Orchestrator

Scalpel v2.0 is an open-source tool that scans your codebase across 12 dimensions and assembles a custom AI surgical team. The entire v2.0 was built in a single Claude Code session using agent teams with worktree isolation.
What Scalpel Does
AI agents are powerful but context-blind - they don't know your architecture, tech debt, git history, or conventions, which leads to guessing and bugs at scale. Scalpel addresses this by:
- Scanning 12 dimensions: stack, architecture, git forensics, database, auth, infrastructure, tests, security, integrations, code quality, performance, documentation
- Producing a Codebase Vitals report with a health score out of 100
- Assembling a custom surgical team where each AI agent owns specific files and gets scored on quality
- Running in parallel with worktree isolation to avoid merge conflicts
Technical Details
The standalone scanner runs in pure bash with zero AI, zero tokens, and zero subscription requirements:
./scanner.sh # Health score in 30 seconds
./scanner.sh --json # Pipe into CI
Sample scans of popular repos:
- Cal.com (35K stars): 62/100 - 467 TODOs, 9 security issues
- shadcn/ui (82K stars): 65/100 - 1,216 'use client' directives
- Excalidraw (93K stars): 77/100 - 95 TODOs, 2 security issues
- create-t3-app (26K stars): 70/100 - zero test files (CRITICAL)
- Hono (22K stars): 76/100 - 9 security issues
Integration and Usage
Scalpel works with 7 AI agents: Claude Code, Codex, Gemini, Cursor, Windsurf, Aider, and OpenCode. It auto-detects your agent on install.
For Claude Code specifically, it's built as a Claude Code agent that lives in .claude/agents/ and activates when you say "Hi Scalpel."
It also ships as a GitHub Action to block unhealthy PRs from merging:
- uses: anupmaster/scalpel@v2
with:
fail-below: 60
comment: true
The v2.0 release includes: scanner + agent brain + 6 adapters + GitHub Action + config schema + tests + docs. The project is MIT licensed with no paid tiers.
📖 Read the full source: r/ClaudeAI
👀 See Also
Researcher Builds Veracity-Checking Skill for Claude Code, Finds Hallucinations in Own Documentation
A researcher built a Claude Code skill called /veracity-tweaked-555 that decomposes documents into atomic claims and verifies each via web search using 16 parallel agents across 4 waves. When self-audited, the skill scored 62/100 due to fabricated statistics and inflated claims in its own documentation.

Browser Harness: Giving LLMs raw CDP access to self-correct browser tasks
Browser Harness strips away browser frameworks, giving LLMs direct CDP websocket access and letting them write missing tools mid-task. Demonstrated by self-inventing an upload_file() function.

LAP: 1,500+ API Specs Compiled for LLM Consumption to Reduce Claude Hallucinations
LAP is a tool that compiles 1,500+ real API specifications into a lean format optimized for LLMs, providing verified endpoints and parameters to prevent AI coding agents like Claude from hallucinating incorrect API calls.

Lumyr: Dashboard Generation via Claude with Python and Streamlit Automation
Lumyr is a tool that generates live, shareable dashboards from plain English descriptions using Claude for dashboard generation and automating the Python and Streamlit layer. Users don't need to write Python, open Streamlit, deploy, set up hosting, or manage infrastructure.