Quant requests?

#5
by pathosethoslogos - opened

Do you take requests for quantising models?

Currently, there are only 3 GGUFs for inclusionAI/Ling-3.0-flash-Fin https://huggingface.co/models?other=base_model:quantized:inclusionAI/Ling-3.0-flash-Fin

NVFP4 of it would be amazing!

@pathosethoslogos They do here https://github.com/NVIDIA/Model-Optimizer

Thanks!

I used the Fin model for a bit, but it actually somehow crashed my vLLM and make my machine hang, a couple of times now. I think vLLM needs to be patched or something, as of today anyway.

Sign up or log in to comment