Issue: CUDA Illegal Instruction on GB10 (Blackwell) with High Context/RAG
#1
by Stor11 - opened
Hi! Thanks for providing this image.
I am running this image on a GB10 (DGX Spark / Blackwell Architecture) with the cybermotaz/nemotron3-nano-nvfp4-w4a16 model.
The Issue: While simple chat works fine with the recommended settings (VLLM_CUDA_GRAPH_MODE: full_and_piecewise), the container crashes immediately when receiving large Context/RAG inputs (> 20k tokens).
Error Log:
torch.AcceleratorError: CUDA error: an illegal instruction was encountered
...
EngineCore encountered a fatal error.
Do you have any solution to this?
Thanks
Thanks for bringing this up! I have noticed this too, and have been working on a fix. I'll keep you posted.