Quant requests?
#5
by pathosethoslogos - opened
Do you take requests for quantising models?
Currently, there are only 3 GGUFs for inclusionAI/Ling-3.0-flash-Fin https://huggingface.co/models?other=base_model:quantized:inclusionAI/Ling-3.0-flash-Fin
NVFP4 of it would be amazing!
@pathosethoslogos They do here https://github.com/NVIDIA/Model-Optimizer
Thanks!
I used the Fin model for a bit, but it actually somehow crashed my vLLM and make my machine hang, a couple of times now. I think vLLM needs to be patched or something, as of today anyway.