Apple Silicon macOS VMs: 11–16× Faster LLM Inference with Metal Capability Shim
Cua — the team behind the Lume macOS virtualization stack — released a research project that patches Metal capability queries inside macOS VMs, unlocking newer GPU kernels. The result: llama.cpp runs 11–16× faster in a macOS VM on Apple Silicon, nearly matching bare-metal performance.
How it works
Apple's Virtualization.framework presents macOS guests with a paravirtualized GPU. The guest's Metal driver reports a conservative capability profile — in stock Tahoe VMs, it reports an Apple 5-era GPU family, limited threadgroup memory, and no SIMD-group matrix support. llama.cpp sees those limits and picks slower kernels, even though the physical GPU can handle more.
Cua's shim intercepts Metal capability queries in a single guest process and returns upgraded values. That lets llama.cpp select newer Metal kernels without changing the guest OS or hypervisor.
Benchmarks
- TinyLlama 1.1B (M1 Ultra): prompt processing 11.08× faster, token generation 16.36× faster vs stock VM. Prompt processing hit 98% of bare-metal.
- Gemma 4 12B QAT Q4_0 (6.98 GB): 7.20× faster prompts, 14.54× faster generation. Reached 99.59% of bare-metal prompt speed, 94.82% generation.
- Muse Glimmer 30B Q4_K-M GGUF (64 GiB guest, llama.cpp b10359): 7.55× faster prompt processing, 8.87× faster generation on a 512-token prompt.
Why it matters
This isn't a GPU passthrough in the VFIO sense — the host still owns the GPU. But it corrects a capability misreport that forced conservative kernel selection. Tart users have filed a similar issue about graphics and LLM performance in macOS guests.
The shim is process-scoped, so it only affects the target app. Everything is released under the same permissive license as Lume and Cua, with source, build scripts, and benchmark logs included.
Who this is for
Developers running llama.cpp or other Metal-based inference inside macOS VMs on Apple Silicon — especially those building local AI tools or testing against VM images.
📖 Read the full source: HN LLM Tools
👀 See Also

Workflow orchestrator with AI CLI integration for sysadmin tasks
A developer built a file-based workflow orchestrator called 'workflow' that integrates with Claude Code, Codex CLI, and Gemini CLI. It generates, updates, fixes, and refines YAML workflows from natural language descriptions for sysadmin tasks.

Context Routing Layer Reduces Claude Code Token Usage by Tracking Accessed Files
A developer saved approximately $80 per month on Claude Code usage by adding a context routing layer that prevents the AI from re-reading the same repository files on follow-up turns. The tool tracks what files have already been accessed to reduce redundant token consumption.

ClawPort: Open Source Orchestration for AI Agent Workflows with Self-Healing Cron
ClawPort is an open source orchestration layer for AI agent workflows that auto-configures cron pipelines, self-heals on failures, and lets you test agents directly before they run on schedule.

Open-Source Claude IDE Bridge Connects Dispatch, Desktop App, and Claude Code
The claude-ide-bridge is an MIT-licensed open-source tool that connects Claude Code to your IDE, providing access to LSP, debugger, terminals, git, and GitHub through 124 tools. It enables a workflow where tasks sent via Dispatch from a phone are handled by the Claude desktop app, which uses Claude Code to write code and run tests while interacting with the IDE.