EvalShift: Open-source CLI for detecting LLM regressions during model migration

EvalShift is an open-source Python CLI designed to detect regressions when switching between LLMs or model versions. It runs your golden input suite against both source and target models, evaluates outputs, and produces a local HTML report — no backend, accounts, or telemetry.
Key features
- Source vs target model comparison via LiteLLM
- JSONL golden suites with tags/slices
- Structural evaluators: JSON schema, regex, length
- Semantic evaluator: embedding similarity
- LLM-as-judge pairwise evaluation
- Tool-call evaluators: tool selection, argument matching, trace structure
- Paired statistical tests: t-test / Wilcoxon
- Effect sizes: Cohen's d
- Multiple-comparison correction: Benjamini-Hochberg
- Slice-level breakdowns
- Local caching to control cost
- Resumable runs
- Single-file HTML report + JSON output
The project's narrow goal is migration safety: “Can I switch models without breaking my prompt/agent behavior?” The author emphasizes catching silent agent regressions — e.g., a newer model producing a decent-looking final answer but skipping a required tool call, calling the wrong tool, or mutating arguments.
Use cases
- Claude 4.5 → Claude 5
- GPT-5 → GPT-6
- Gemini 2 → 3
- Local model → hosted model
The author is seeking feedback on usefulness for local vs hosted models, most important evaluator types for local LLM workflows, and whether tool-call/structured-output regressions are a real pain point. The repo is MIT licensed.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Strale.io offers free IBAN and email validation API for AI agents with no signup
Strale.io provides a free API with five capabilities including IBAN validation, email validation, DNS lookup, URL-to-markdown conversion, and JSON repair. No signup or API key is required, and it includes an MCP server for Claude or Cursor integration.

Lore: A tool that extracts structured context from AI coding conversations
Lore is a browser-based tool built with Claude Code that extracts structured context from AI conversations, capturing decisions, TODOs, blockers, and resume checklists. It's a React + TypeScript PWA with a Chrome extension for direct conversation capture and context injection.

Focusmo macOS app adds local MCP server for Claude AI integration
Focusmo, a macOS focus app, now includes a local MCP server that allows Claude AI to access real focus data for weekly reviews and planning. The server runs locally on Mac with no external servers required, keeping all data on-device.
Needle: A 26M Parameter Tool-Calling Model Built Entirely Without FFNs
Needle is a 26M parameter function-calling model with no MLPs, achieving 6000 tok/s prefill and 1200 tok/s decode on consumer devices. It beats FunctionGemma-270M, Qwen-0.6B, Granite-350M, and LFM2.5-350M on single-shot tool calling.