PhAIL Benchmark Tests VLA Models on Real Warehouse Robot Tasks

PhAIL is a physical AI benchmark that measures how well vision-language-action (VLA) models perform on commercial robotics tasks. The creator built it because they couldn't find honest performance numbers for these models in practical applications.
Benchmark Details
The benchmark tests four VLA models on bin-to-bin order picking, one of the most common warehouse operations:
- OpenPI/pi0.5
- GR00T
- ACT
- SmolVLA
All tests use the same equipment: a Franka FR3 robot with Robotiq 2F-85 gripper (DROID setup), with identical objects across hundreds of blind runs where the operator doesn't know which model is running.
Performance Results
The benchmark revealed significant performance gaps:
- Best model performance: 64 units per hour (UPH)
- Human teleoperating the same robot: 330 UPH
- Human performing the task by hand: 1,300+ UPH
Open Data and Methodology
Everything from the benchmark is publicly available:
- Every run with synced video and telemetry data
- The fine-tuning dataset used for training
- Training scripts
- An open leaderboard accepting new submissions
The creator is available to answer questions about methodology, the specific models tested, or observations from the benchmark runs.
📖 Read the full source: HN AI Agents
👀 See Also

OpenClaw Implements Agent History Compression to Reduce Context Usage
OpenClaw now compresses agent history by replacing completed subtask logs with structured summaries, reducing ~1M tokens to ~30K. The system uses a 4-pass scanner to identify task lifecycles and generates masked summaries that maintain agent compatibility.

mistral.rs Adds Support for Gemma 4 12B: Multimodal, Agentic, and MTP
mistral.rs now supports Gemma 4 12B with multimodal, agentic, and MTP integration. One-step install and run with web search, code execution, and built-in UI.

Codev: AI agent workflow for 106 PRs in 14 days
Codev is an open-source system that coordinates multiple AI agents through a strict Spec→Plan→Implement→Review→PR workflow, catching 20 bugs before shipping and producing code rated 1.2 points better on a 10-point scale.

Claude's Code Dashboard Tracks 19M+ AI-Generated Commits on GitHub
A developer built a dashboard tracking over 19 million commits generated by Claude Code on GitHub public repositories, showing TypeScript (35.3%), Python (19.2%), and JavaScript (10.3%) as the top languages. The system uses Next.js with Recharts and PostgreSQL, with an ETL pipeline that works around GitHub's API rate limits.