Running LLMs on NVIDIA Blackwell: vLLM, SGLang, and NVFP4

On Blackwell GPUs, vLLM with NVFP4 quantization gave the highest batched throughput we measured: 8,430 tokens per second on Llama 3.1 8B on one RTX PRO 6000. SGLang beat vLLM by 17% on identical GPTQ models and is what we used to fit the largest models onto a single card. Ollama is for development, not serving. The hard part is rarely the model. It's pinning a working engine build, choosing a parallelism strategy that fits your interconnect, and fitting the weights and KV cache into VRAM.

Everything here was measured on consumer and workstation Blackwell cards in the Joshua8.AI lab. Each post includes the exact image, flags, and results.

Best results by card

CardResultSource
RTX 5060 Ti 16 GB1,259 tok/s on Llama 3.1 8BGPU shootout
2× RTX 5070 Ti (no P2P)Pipeline parallel ran prefill 1.9× faster than tensor parallel on Qwen3.6-35BTwo 5070 Tis, Part 1
RTX 5090 32 GBGLM-4.7-Flash at full 128K context, over 100 tok/sGLM-4.7-Flash setup
RTX PRO 6000 96 GB8,430 tok/s batched on Llama 3.1 8B (vLLM, NVFP4)Engine benchmark update
RTX PRO 6000 96 GBMiniMax-M2.5 (139B MoE) at 68 tok/s with 64K contextMiniMax-M2.5 guide
RTX PRO 6000 96 GB127 GB Qwen3.8-Flash-Next NVFP4 at 256K context, 94 tok/sQwen3.8-Flash-Next guide

Choosing and pinning an engine

Bring-up guides by card

Quantization and precision

Throughput and concurrency

Related

Bringing up a model on Blackwell hardware, or deciding which card to buy? We do this every week.

Talk to Us