Running LLMs on NVIDIA Blackwell: vLLM, SGLang, and NVFP4
On Blackwell GPUs, vLLM with NVFP4 quantization gave the highest batched throughput we measured: 8,430 tokens per second on Llama 3.1 8B on one RTX PRO 6000. SGLang beat vLLM by 17% on identical GPTQ models and is what we used to fit the largest models onto a single card. Ollama is for development, not serving. The hard part is rarely the model. It's pinning a working engine build, choosing a parallelism strategy that fits your interconnect, and fitting the weights and KV cache into VRAM.
Everything here was measured on consumer and workstation Blackwell cards in the Joshua8.AI lab. Each post includes the exact image, flags, and results.
Best results by card
| Card | Result | Source |
|---|---|---|
| RTX 5060 Ti 16 GB | 1,259 tok/s on Llama 3.1 8B | GPU shootout |
| 2× RTX 5070 Ti (no P2P) | Pipeline parallel ran prefill 1.9× faster than tensor parallel on Qwen3.6-35B | Two 5070 Tis, Part 1 |
| RTX 5090 32 GB | GLM-4.7-Flash at full 128K context, over 100 tok/s | GLM-4.7-Flash setup |
| RTX PRO 6000 96 GB | 8,430 tok/s batched on Llama 3.1 8B (vLLM, NVFP4) | Engine benchmark update |
| RTX PRO 6000 96 GB | MiniMax-M2.5 (139B MoE) at 68 tok/s with 64K context | MiniMax-M2.5 guide |
| RTX PRO 6000 96 GB | 127 GB Qwen3.8-Flash-Next NVFP4 at 256K context, 94 tok/s | Qwen3.8-Flash-Next guide |
Choosing and pinning an engine
- Benchmarking LLM Inference: vLLM vs SGLang vs Ollama on Blackwell vLLM NVFP4 hit 8,033 tok/s; SGLang beat vLLM by 17% on GPTQ; Ollama is 10× slower.
- Do LLM Inference Engines Actually Get Faster With Updates? A month later, every engine improved, but fixing Ollama's config beat any version bump.
- vLLM v0.16.0: Up to 20% Higher Throughput on Blackwell 8–10% at moderate concurrency and 20% at 128 concurrent requests.
- Chasing an 8% Decode Regression in vLLM Nightlies Bisecting 315 commits to an 8-line PR that traded throughput for compile speed.
- The Nightly From Hell: Bisecting vLLM for a Working NVFP4 Build An undocumented SM120 crash fix and nine nightlies to find a clean pin.
Bring-up guides by card
- Running Qwen3.5-35B on an RTX 5090 with vLLM Fixing the Triton autotuner OOM with a cross-GPU cache warmup.
- GLM-4.7-Flash: 128K Context on a Single Consumer GPU MLA and vLLM settings for full context on 32 GB.
- MiniMax-M2.5 on a Single RTX PRO 6000 A 139B MoE at 68 tok/s with SGLang and NVFP4.
- A 127 GB Model on a 96 GB Card Offloading the n-gram table to host RAM, the pinned SGLang image, and the one patch you still need.
- Two RTX 5070 Tis: Pipeline vs Tensor Parallel Without a P2P link, the interconnect dictates the split.
- Evicting the Vision Encoder to the CPU A CPU sidecar that freed VRAM, and why the boring config still won production.
Quantization and precision
- GLM-4.7-Flash: NVFP4 vs AWQ AWQ runs 6–17% faster; NVFP4 holds a slight accuracy edge.
- When FP16 Is Faster Than Quantized Quantized models can take longer to answer because they generate more thinking tokens.
- The Quantization Ladder Has Broken Rungs Q6_K beat Q8_0 on accuracy; the smooth quantization curve is a myth.
Throughput and concurrency
- Scaling vLLM: How Parallel Queries Change Everything Near-linear scaling with no accuracy loss: 9.2× faster at 16 concurrent queries.
- The Decision Chart: Which Model Should You Actually Use? What to deploy by VRAM, task, and latency after 300,000+ inference calls.
- Why Speculative Decoding Didn't Speed Up a 120B Model Eagle3 on GPT-OSS-120B on one RTX PRO 6000: baseline won.
Related
- Running AI on eBay Hardware The other end of the lab: 200B-class models on a $1,000 Xeon with 2016-era Pascal GPUs.
- Local AI Economics: Own the Hardware or Rent the Cloud? Break-even math, local RAG speedups, and what GPUs really cost to own and run.
- Our Companies The same lab builds TeraContext.AI, our AI document intelligence company for commercial construction.
Bringing up a model on Blackwell hardware, or deciding which card to buy? We do this every week.
Talk to Us