Agent-Xray: Open-source tool for debugging AI agent failures from trace logs

Agent-Xray is an open-source tool for debugging AI agents by analyzing their trace logs. It was created to solve the problem of agents failing tasks without clear errors—situations where code runs fine but the agent makes wrong decisions, like repeatedly calling the wrong tool despite error messages suggesting the correct one.
Key Features
The tool reads trace logs and provides structural grading and root-cause classification for agent failures. It reconstructs what the agent was seeing at each step to help understand why bad decisions were made.
Failure Categories
- spin
- tool_bug
- early_abort
Enforcement Mode
The most significant feature according to the creator is enforcement mode. After fixing an agent bug, this mode runs adversarial challenges against your fixes to verify they're legitimate. It checks for:
- Hardcoded returns
- Weakened assertions
This addresses the problem where fixes might work on specific test tasks but are actually fragile, or where agents learn to game the test.
Workflow Integration
The tool runs as MCP tools, allowing Claude Code to use it directly. A typical workflow described in the source:
- Tell Claude Code to triage agent traces
- It finds the worst failure
- Replays what the agent saw
- Suggests a fix
- Enforcement mode verifies the fix is legitimate
The creator describes this as "agents debugging agents."
Technical Details
- Installation:
pip install agent-xray - Quickstart:
agent-xray quickstart(includes sample traces to test without your own data) - License: MIT
- Zero dependencies
- Runs offline
- Works with OpenAI, Anthropic, LangChain, CrewAI, OpenTelemetry traces
- Project age: About 9 days old at time of posting
Use Case
This tool is for developers working with AI agents who need to debug failures that don't produce traditional errors or stack traces—situations where agents make incorrect decisions despite having access to correct tools and information.
📖 Read the full source: r/ClaudeAI
👀 See Also

Definable AI adds self-hosted observability dashboard with single flag
Definable AI, an open-source Python framework for building AI agents, now includes a built-in observability dashboard that can be enabled with one flag. The dashboard provides real-time event streaming, token accounting, latency metrics, and run replay without external dependencies.

Femtobot: Efficient Rust Agent for Low-Resource Environments
Femtobot is a lightweight Rust-based AI agent designed to run efficiently on low-resource machines, such as older Raspberry Pis, through a ~10MB binary without large runtime dependencies.

ByteRover Memory Plugin for OpenClaw: Native Integration with Semantic Hierarchy
ByteRover Memory Plugin for OpenClaw provides native, structured long-term memory via a three-layer architecture and semantic hierarchy stored in Markdown files. It achieves 92.2% retrieval accuracy and requires OpenClaw v2026.3.22+.

Headless OpenClaw Setup with Discord via Docker Scripts
A GitHub repository provides scripts to run OpenClaw with Discord in a headless Docker container, avoiding the TUI/WebUI. It includes a management script with commands like claw init, start, and stop, plus preconfigured support for OpenAI Responses API, Chromium, and various tools.