Instructions to use unsloth/Qwen3.8-27B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
TP2 benchmark on 2x RTX 5070 Ti (16GB) -- vision works, decode speed + VRAM findings
Benchmark: 2x RTX 5070 Ti (16GB each, sm_120) TP2, vLLM cu129-nightly
Vision: WORKS -- correctly identified objects in test images. Vision encoder intact.
Config: TP2, GMU 0.95, FP8 KV, 8GB CPU KV offload, MTP K=3, flashinfer, compressed-tensors, max-model-len 80K
VRAM per GPU (16GB): weights 11.3 GiB, KV 1.51 GiB (80K tokens), CPU offload 4 GiB (~200K tokens), CUDA graphs 0.14 GiB, activation 1.43 GiB. Total ~14.4/15.47 GiB.
Decode (MTP K=3, accept_len ~2.2): 54 tok/s at ctx=0, 64 tok/s at ctx=64K. Prefill: 2.7-3.8K tok/s.
vs modelopt NVFP4 (pottokao checkpoint): Unsloth is 1.4 GiB/card heavier (11.3 vs 9.9), reducing KV headroom (80K vs 120K context) and decode speed (54-64 vs ~120 tok/s). But vision works on unsloth while pottokao crashes on TP2. For 24GB+ cards the weight difference matters less.
Thanks for the day-zero NVFP4 release!
nice - I would like to know your optimal performance setup - particularly with Unsloth Dynamic V3.0 (preview) - My second 5070 ti is coming in the mail this week. Excited to get this all set up!
Hello is it possible to give us your command line please ? I’ve the same setup as yours.
Thx !
Benchmark: 2x RTX 5070 Ti (16GB each, sm_120) TP2, vLLM cu129-nightly
Vision: WORKS -- correctly identified objects in test images. Vision encoder intact.
Config: TP2, GMU 0.95, FP8 KV, 8GB CPU KV offload, MTP K=3, flashinfer, compressed-tensors, max-model-len 80K
VRAM per GPU (16GB): weights 11.3 GiB, KV 1.51 GiB (80K tokens), CPU offload 4 GiB (~200K tokens), CUDA graphs 0.14 GiB, activation 1.43 GiB. Total ~14.4/15.47 GiB.
Decode (MTP K=3, accept_len ~2.2): 54 tok/s at ctx=0, 64 tok/s at ctx=64K. Prefill: 2.7-3.8K tok/s.
vs modelopt NVFP4 (pottokao checkpoint): Unsloth is 1.4 GiB/card heavier (11.3 vs 9.9), reducing KV headroom (80K vs 120K context) and decode speed (54-64 vs ~120 tok/s). But vision works on unsloth while pottokao crashes on TP2. For 24GB+ cards the weight difference matters less.
Thanks for the day-zero NVFP4 release!