vLLM Setup and Testing on 10x NVIDIA V100 Server with 320GB VRAM

Hardware Configuration and Build Notes
A developer has built a local AI server with 10x Tesla V100 SXM2 32GB GPUs (320GB VRAM total) on an AMD Threadripper PRO system. The setup uses Ubuntu 24.04 headless with NVIDIA driver 580.126.20. GPU topology consists of two NVLink quad meshes (GPUs 0-3, 4/5/8/9) plus an NV6 pair (GPUs 6-7).
What Works on V100 with vLLM
- FP16 unquantized: Primary path using
--dtype half - bitsandbytes 4-bit: Works for models too large for FP16
- TRITON_ATTN: Automatic fallback since FlashAttention2 requires SM 80+
- Tensor/Pipeline parallel: TP=4 and TP=4 PP=2 both tested successfully
What Does Not Work on V100
- GPTQ: ExLlamaV2 kernels broken on SM 7.0 (vLLM issue #2165)
- AWQ: Requires SM 75+
- FP8: Requires SM 75+. MiniMax M2.5 uses FP8 internally — dead on arrival.
- FlashAttention2: Requires SM 80+
- DeepSeek MLA: Hopper/Blackwell only. Full DeepSeek V3/R1 cannot run on vLLM + V100.
Build Requirements and Critical Fixes
PyTorch 2.11.0+cu126 is required — cu126 is the last version with V100 support as cu128+ drops Volta. Source compilation requires TORCH_CUDA_ARCH_LIST="7.0" and MAX_JOBS=20. A MoE kernel patch is needed for issue #36008, changing B.size(1) to B.size(0) in fused_moe.py (2 lines). PYTHONNOUSERSITE=1 is required to isolate conda environment from stale system packages.
Critical NCCL Dependency Fix: pip install -e . pulls in nvidia-nccl-cu13 alongside nvidia-nccl-cu12. The cu13 library gets loaded at runtime and references CUDA 13 symbols that don't exist in the cu126 runtime, resulting in "NCCL error: unhandled cuda error" on every multi-GPU launch. The fix involves uninstalling all nvidia-* packages and managing dependencies carefully.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Fine-Tuning Qwen 3:0.6B for Question Categorization – Baseline vs Finetuned Results
Fine-tuning a tiny 0.6B parameter LLM (Qwen 3:0.6B) with ~850 household questions using Unsloth. Baseline prompting scored 10% accuracy; finetuned results likely exceed 80-90%.

System Architecture for Vibe Coders: A Senior Engineer's Guide
A 10-year engineer shares how to approach app building with Claude Code: start at the system level, not the code. Covers the four components — frontend, backend, database, plumbing — with a deep dive into the plumbing: APIs, hosting, deployment, secrets, and security.

Todoist connector removed from Claude, custom setup required
The official Todoist connector is no longer available in Claude. Users can add Todoist as a custom connector using the MCP URL https://ai.todoist.net/mcp, but this requires a Claude Pro or Max subscription.

iOS Shortcut Workaround for Sending iPhone Photos to Cowork via iCloud Sync
A developer created an iOS Shortcut called "PhoPo" that converts iPhone photos to JPEG, resizes them, and saves them to an iCloud-synced folder that Cowork can access, enabling Claude to analyze screenshots and photos from mobile devices.