MTP + Unified Memory Boosts llama.cpp Inference 30% on RTX 5090

✍️ OpenClawRadar📅 Published: May 12, 2026🔗 Source
Ad

Combining GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 with Multi-Token Prediction (MTP) speculation in llama.cpp yields a ~30% throughput improvement — 64 tok/sec vs 49 tok/sec on a Qwen3.6-27B Q8_0 model. The benchmark was run on an RTX 5090 paired with 128GB DDR5 5600 CL36 and a Ryzen 9 9950X3D.

Command & Configuration

CUDA_VISIBLE_DEVICES=0 GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 /home/marcin/llama-server \
    -m /home/marcin/Pobrane/Qwen3.6-27B-Q8_0.gguf \
    --threads 16 \
    -c 262144 -fa on -np 1 \
    --spec-type mtp --spec-draft-n-max 3 \
    --webui-mcp-proxy \
    --chat-template-kwargs '{"preserve_thinking": true}' \
    --host 0.0.0.0 \
    --port 8090 \
    --jinja

Key flags:

  • GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 — allows the GPU to directly access host memory, bypassing CUDA malloc for large contexts.
  • --spec-type mtp --spec-draft-n-max 3 — enables Multi-Token Prediction speculation with a draft depth of 3.
  • Qwen3.6-27B-Q8_0.gguf — a 27B parameter Qwen3.6 model quantized to Q8_0, prepared with Unsloth’s MTP support.
  • -c 262144 — 256K context window; -fa on for flash attention.
Ad

Results

  • Without MTP (only unified memory): 49 tok/sec
  • With MTP + unified memory: 64 tok/sec
  • Gain: 30% higher throughput

The draft-n-max of 3 means the model speculates up to 3 tokens ahead, reducing serial decoding overhead. Combined with unified memory, it avoids expensive PCIe transfers between CPU and GPU RAM.

Who This Is For

Developers running large-context local inference on high-end consumer GPUs (RTX 5090) with ample system RAM (≥128GB). Suitable for chatbots, code assistants, or any latency-sensitive LLM workload where speculative sampling is supported.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

ETL-D MCP Server: Deterministic CSV Parsing for Claude to Prevent Financial Hallucinations
Tools

ETL-D MCP Server: Deterministic CSV Parsing for Claude to Prevent Financial Hallucinations

A developer built ETL-D, an open-source MCP server for Claude Desktop that processes CSVs in three deterministic layers to prevent decimal point hallucinations in financial data. It uses Python parsers for known formats, achieves ~70ms response times with 0 LLM calls for 200 parallel requests, and only uses LLMs as a fallback for high-entropy text.

OpenClawRadar
Claw Code Agent: Python Reimplementation of Claude Code Architecture for Local Models
Tools

Claw Code Agent: Python Reimplementation of Claude Code Architecture for Local Models

Claw Code Agent is a Python reimplementation of the Claude Code agent architecture that runs with local open-source models through OpenAI-compatible backends like vLLM and Ollama, featuring tool calling, slash commands, and tiered permissions.

OpenClawRadar
Sonicker: Voice Cloning Web App Built with Claude Code in 4 Days
Tools

Sonicker: Voice Cloning Web App Built with Claude Code in 4 Days

Sonicker is a voice cloning web app that requires only 3 seconds of audio input and supports 10 languages. The developer built it solo in 4 days using Claude Code for the entire frontend, API integration, and deployment.

OpenClawRadar
Developer Builds Tool for Realistic Relational Database Generation
Tools

Developer Builds Tool for Realistic Relational Database Generation

A developer built a tool that generates fully loaded relational databases with realistic data, solving the problem of creating test databases with intact foreign key relationships and cross-table consistency.

OpenClawRadar