Issue: CUDA Illegal Instruction on GB10 (Blackwell) with High Context/RAG

#1
by Stor11 - opened

Hi! Thanks for providing this image.

I am running this image on a GB10 (DGX Spark / Blackwell Architecture) with the cybermotaz/nemotron3-nano-nvfp4-w4a16 model.

The Issue: While simple chat works fine with the recommended settings (VLLM_CUDA_GRAPH_MODE: full_and_piecewise), the container crashes immediately when receiving large Context/RAG inputs (> 20k tokens).

Error Log:

torch.AcceleratorError: CUDA error: an illegal instruction was encountered
...
EngineCore encountered a fatal error.

Do you have any solution to this?
Thanks

Thanks for bringing this up! I have noticed this too, and have been working on a fix. I'll keep you posted.

Sign up or log in to comment