Running a 1 Trillion Parameter LLM Locally on AMD Ryzen AI Max+ Cluster

✍️ OpenClawRadar📅 Published: March 1, 2026🔗 Source
Running a 1 Trillion Parameter LLM Locally on AMD Ryzen AI Max+ Cluster
Ad

Running a 1 Trillion Parameter LLM Locally on AMD Ryzen AI Max+ Cluster

AMD's technical article details how to build a small-scale distributed inference cluster using four Framework Desktop systems with Ryzen AI Max+ 395 processors and run the Kimi K2.5 open-source model (1 trillion parameters, 375GB) using llama.cpp RPC. The setup treats the four machines as a single logical AI accelerator.

Hardware and Software Stack

  • Hardware: 4x Framework Desktop - AMD Ryzen AI Max+ 395 - 128GB
  • AI Framework: AMD ROCm
  • Inference Engine: Llama.cpp RPC
  • OS: Ubuntu 24.04.3 LTS
  • Model: Kimi-K2.5 (UD_Q2_K_XL) (375GB)
  • Network: 5Gbps over Ethernet

Technical Setup: Extended VRAM Allocation

For each Ryzen AI Max+ system, BIOS must first set iGPU Memory Size to 512MB. The maximum dedicated VRAM per node via BIOS is 96GB (384GB total across four nodes). Using Translation Table Manager (TTM) kernel parameters increases this to 120GB per node (480GB total).

Configure kernel parameters:

sudo nano /etc/default/grub

Find line starting with GRUB_CMDLINE_LINUX_DEFAULT= and append inside quotes:

"quiet splash ttm.pages_limit=30720000 amdgpu.gttsize=120000"

TTM limits are expressed in 4 KB pages. Calculation for 120GB: (120 * 1024 * 1024) / 4.096 = 30720000

After saving and exiting, run:

sudo update-grub
sudo reboot

Verify configuration:

$ sudo dmesg | grep "amdgpu.*memory"
[drm] amdgpu: 512M of VRAM memory ready
[drm] amdgpu: 120000M of GTT memory ready.
Ad

Setup Option 1: Lemonade SDK (Recommended)

Download pre-built binaries from: https://github.com/lemonade-sdk/llamacpp-rocm/releases/latest/

Download archive matching your platform and GPU target: llama-bxxxx-ubuntu-rocm-gfx1151-x64.zip

Extract and prepare:

unzip llama-bxxxx-ubuntu-rocm-gfx1151-x64.zip
cd llama-bxxxx-ubuntu-rocm-gfx1151-x64
chmod +x llama-cli llama-server rpc-server

Verify GPU detection:

$ ./llama-cli --list-devices
ggml_cuda_init: found 1 ROCm devices:
Device 0: AMD Radeon Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32
Available devices:
ggml_backend_cuda_get_available_uma_memory: final available_memory_kb: 127697544
ROCm0: AMD Radeon Graphics (120000 MiB, 124704 MiB free)

Setup Option 2: Manual Source Build

Install ROCm 7.0.2 on Ubuntu 24.04.3:

wget https://repo.radeon.com/amdgpu-install/7.0.2/ubuntu/noble/amdgpu-install_7.0.2.70002-1_all.deb
sudo apt install ./amdgpu-install_7.0.2.70002-1_all.deb
sudo apt update
sudo apt install python3-setuptools python3-wheel
sudo usermod -a -G render,

The article continues with additional setup steps and inference configuration details.

📖 Read the full source: HN LLM Tools

Ad

👀 See Also

Practical Review: 3 Essential Clawhub Skills and 3 to Avoid
Guides

Practical Review: 3 Essential Clawhub Skills and 3 to Avoid

A developer tested Clawhub skills for weeks and found three worth installing: web-search (Brave), daily-brief, and memory-search. Three others—food-order, multi-agent orchestrators, and humanizer—waste tokens and add unnecessary complexity.

OpenClawRadar
OpenClaw Failure Patterns: 42 Real Incidents in 28 Days
Guides

OpenClaw Failure Patterns: 42 Real Incidents in 28 Days

A developer running OpenClaw daily documented 42 specific failures across eight categories, including AI hallucinations, authentication breakdowns, and automation that costs more time than it saves. The source provides concrete examples like Google OAuth 7-day token expiration and Opus 4.6 adding unwanted metadata to files.

OpenClawRadar
OpenClaw setup guide from Reddit analysis: hardware, cost, memory, and security practices
Guides

OpenClaw setup guide from Reddit analysis: hardware, cost, memory, and security practices

A Reddit user analyzed common OpenClaw mistakes and created a setup guide covering hardware requirements, cost optimization to $10/month, memory management using MEMORY.md files, and security practices to prevent prompt injection attacks.

OpenClawRadar
Developer shares 25 tested Claude prompts for SaaS development workflows
Guides

Developer shares 25 tested Claude prompts for SaaS development workflows

A developer has shared 25 specific prompts they use daily for SaaS development, covering backend architecture, API design, frontend copy, product documentation, and go-to-market tasks. The prompts are designed to save time on repetitive tasks like code review, documentation generation, and edge case testing.

OpenClawRadar