Performance Tip: Lock Local Model VRAM/RAM with LimitMEMLOCK=infinity
If your local model fits inside your combined VRAM and RAM, you can keep the model provider's memory resident by adding LimitMEMLOCK=infinity under the [Service] header of the systemd unit for your provider. The kernel then stops paging model weights out to disk, which is where most of the slowdowns come from once a model is loaded.
The unit file
Posted in r/openclaw as a working example for LM Studio (edit the paths and usernames to match your install):
[Unit] Description=LM Studio Server RequiresMountsFor=/home[Service] LimitMEMLOCK=infinity Type=oneshot RemainAfterExit=yes User=******** Environment="HOME=/home/" ExecStartPre=/home//.lmstudio/bin/lms daemon up ExecStartPre=/home//.lmstudio/bin/lms load text-embedding-bge-large-en-v1.5 --yes ExecStartPre=/home//.lmstudio/bin/lms load qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive --yes ExecStart=/home//.lmstudio/bin/lms server start ExecStop=/home//.lmstudio/bin/lms daemon down
[Install] WantedBy=multi-user.target
What each part does
LimitMEMLOCK=infinity— removes the default per-process mlock cap, so the loaded model can be kept out of swap/paging.ExecStartPrewithlms daemon up— starts the LM Studio daemon before anything else.lms load <model> --yes— preloads each model at service start. The example loads an embedding model (text-embedding-bge-large-en-v1.5) and a chat model (qwen3.6-35b-a3b-uncensored-hauhaucs-aggressive).lms server startasExecStart— brings up the OpenAI-compatible server.lms daemon downasExecStop— cleans up on shutdown.RequiresMountsFor=/home— ensures the home partition holding the model files is mounted before the service starts.Type=oneshot+RemainAfterExit=yes— the unit is treated as active after the start commands finish, which fits a daemon thelmsCLI manages separately.
Caveat
The tip explicitly assumes the model fits in VRAM+RAM. If it doesn't, locking memory is the wrong lever — you'll just push the problem elsewhere. Once it fits, the win is avoiding paging-induced stalls on inference.
To apply: drop the block into your unit file, then systemctl daemon-reload and restart the service. Developers running LM Studio headless as a systemd service are the target audience here; the same idea applies to any provider process whose weights you want pinned in memory.
📖 Read the full source: r/openclaw
👀 See Also

Claude's Research Output Varies by Language: Same Prompt, Different Sources
A Reddit test shows Claude returning different sources and developments across English, Chinese, Russian, Spanish, and Hindi prompts — same model, same structure, diverging results.

How to Prevent CLAUDE.md Rot: Treat Rules Like Code
After 18 months of real-world use, one developer shares four disciplines to keep CLAUDE.md under 100 lines: use it as an index, separate rules from sources, audit on every PR, and delete more than you add.

Essential Custom Instructions for Claude to Prevent Common Annoyances
A Reddit user shares three specific custom instructions to address common Claude annoyances: requiring warnings before destructive commands, preventing mid-answer plan changes, and keeping code blocks exclusively for functional code.

AGENTS.md Pattern for React Native: Claude Code Generates Better Project-Aware Code
A Reddit user shares their AGENTS.md file for React Native/Expo projects that includes folder structure, theme tokens, custom hooks, and component patterns. The result: Claude Code and Cursor generate code using the exact project conventions instead of generic React Native code.