Run OpenClaw with a Local LLM on macOS – Guide for 16–24GB RAM

A new guide walks through setting up OpenClaw with a local LLM on macOS, specifically targeting machines with 16–24GB RAM. The author tested a quantized version of Qwen 3.5 configured for OpenClaw, and includes a test skill to confirm everything is working.
Setup Overview
- Model: Qwen 3.5 (quantized) – chosen to fit within 16–24GB RAM while providing decent reasoning capability.
- Platform: macOS (tested on Mac Mini with 16–24GB).
- Key step: Configure OpenClaw to use the local model endpoint (typically via Ollama or llama.cpp). The guide provides specific config file edits.
Test Skill
To validate the setup, the author created a test skill that calls the local model and returns a known response. If the skill executes correctly, your local LLM is fully integrated with OpenClaw.
Why Local LLM?
Running an LLM locally avoids API costs and latency, keeps code and prompts on-device, and works offline. For OpenClaw users with Apple Silicon Macs, quantized models like Qwen 3.5 are a practical compromise between accuracy and memory.
Next Steps
If the test skill fails, check your model server (Ollama) is running and the OpenClaw config points to the correct URL (http://localhost:11434 for Ollama). Adjust context window size if needed to fit memory.
📖 Read the full source: r/openclaw
👀 See Also

OpenClaw 5.28: Codex Plugin Broken After Upgrade — Fix with Symlink Shim
OpenClaw 5.28 breaks Codex plugin due to binary path mismatch. Fix: create symlink from expected path to actual bin/codex.

Qwen3.5-397B MoE Runs on 14GB RAM via Paged Expert Loading on M1 Ultra
Paged MoE engine keeps only 20 experts resident and lazy-loads the rest from SSD, running a 209GB 397B model on a 64GB Mac Studio with 1.59 tok/s and 14GB peak RAM. Includes smaller model benchmarks.

Fixing Claude Code's KV Cache Invalidation with Local Backends
Claude Code versions 2.1.36+ inject dynamic telemetry headers and git status updates into every request, breaking prefix matching and forcing full 20K+ token system prompt reprocessing on local backends like llama.cpp. A configuration fix in ~/.claude/settings.json can reduce processing from 60+ seconds to ~4 seconds.

5 Common OpenClaw Setup Mistakes and How to Fix Them
Practical fixes for the five most common OpenClaw setup mistakes: skipping persistent memory, no outbound access, overloading system prompt, missing fallback behavior, and using a single model.