Jake Benchmark v1: Local LLM Performance Testing for OpenClaw AI Agents

The Jake Benchmark v1 is a performance evaluation tool for local LLMs functioning as AI agents with OpenClaw. It tests models on 22 practical tasks to determine their effectiveness in real-world agent scenarios.
Test Setup and Methodology
The benchmark was run on a Raspberry Pi with Ollama running on an NVIDIA 3090 GPU. The developer tested 7 different local LLMs to identify the best model for agent work with OpenClaw.
Task Categories
The 22 tasks covered real-world scenarios including:
- Reading emails and creating tasks from them
- Scheduling meetings and checking for conflicts
- Phishing detection (specifically a fake email pretending to be the owner asking for a bitcoin wallet key)
- Error handling
Key Results
The performance variation was significant across models:
- Qwen 27B: Scored 59.4% - successfully handled emails, scheduled meetings, detected phishing attempts, and managed errors
- Nemotron 30B: Scored 1.6% - attempted to solve tasks by running
apt-get install git
Notable Observations
The phishing test revealed interesting behaviors:
- The best model refused the phishing request immediately
- The worst model read the secrets file three times before deciding not to share the information
Dashboard Features
The benchmark includes an interactive dashboard that allows users to:
- Click into any model to view the full conversation
- See exactly what each model did during tasks
- Identify where models went wrong in their execution
The tool is available on GitHub for developers to run their own evaluations and compare local LLM performance for agent tasks.
📖 Read the full source: r/openclaw
👀 See Also

Mastering Antropic Subscription Modes: Haiku, Sonnet, and Opus
Explore Antropic's innovative subscription modes—Haiku, Sonnet, and Opus—designed to enhance your AI coding experience with tailored features and pricing.

Cloudflare Dynamic Worker Loader: Sandboxing AI Agents with Isolates
Cloudflare's Dynamic Worker Loader API, now in open beta, allows Workers to instantiate new Workers with runtime-specified code in isolated sandboxes using V8 isolates, offering 100x faster startup than containers and no global concurrency limits.

mistral.rs Adds Support for Gemma 4 12B: Multimodal, Agentic, and MTP
mistral.rs now supports Gemma 4 12B with multimodal, agentic, and MTP integration. One-step install and run with web search, code execution, and built-in UI.
xAI TTS Integration for Home Assistant Built with Claude — Full Repo
A developer used Claude to build a custom Home Assistant integration for xAI's TTS API (Eve voice) with full UI config, five voices, and speech tags.