No speculative decoding, No QAT

#21
by engrtipusultan - opened

Hi team

Please look at other teams like LFM, Google gemma. Everyone shipping MTP, or Dflash or Dspark for speculative decoding

Also LFM is releasing QAD model that is 97% accurate to BF16 with dspark speculative decoding for 3B model. Similarly gemma has MTP with QAT releases.

I hope you guys take inspiration from that.

The companies you mention actually make LLMs, ornith-ai so far only make finetunes of existing LLMs, they re-trained existing LLMs with their own datasets or whatever, however the MTPs do not work correctly because they were made by Qwen for the oiriginal/official model, not for the onith-ai finetune(s).

MTP heads are not that hard to train (and can and probably should be done during the tuning stage) - luckily it's also easy to do afterwards (mine was done in less than a day: https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY ).

DFlash does takes more effort (days to get a decent version). QAT is also practically free while training, but the QAT is targeted for a single specific quant, so if home users are mostly going to use various GGUF quants, it doesn't much matter.

MTP heads are not that hard to train (and can and probably should be done during the tuning stage) - luckily it's also easy to do afterwards (mine was done in less than a day: https://huggingface.co/shisa-ai/Ornith-1.5-35B-A3B-MTP-ONLY ).

DFlash does takes more effort (days to get a decent version).

True that is the point. They can if they are willing too.

QAT is also practically free while training, but the QAT is targeted for a single specific quant, so if home users are mostly going to use various GGUF quants, it doesn't much matter.

QAT or QAD reduces error when it is quantized to Q4. Which is not same as quants when QAT or QAD is not done. Take LFM example

I think the more important point is QAT has to be targeted for a specific target, eg NVFP4, MXFP4, GPTQ W4A16, Q4_0, etc. This is only useful for someone running an "official" quant for that specific target/layout. In practice, almost all "local" users are running GGUF quants. A Q4_K is actually a mix of Q4_K, Q5_K, Q6_K and Q8_0 (and some F32) tensors.

Note: as Unsloth showed w/ the Gemma QAT models even a Q4_0 safetensor -> Q4_0 GGUF does not protect accuracy/causes severe degradation unless you do additional work.

Gemma team has provided Q4_0 and same for LFM Q4_0. Q4_0 is easy inference for all backends. Only gpt oss was MXFP4. That is also properly supported on all backends in llama.cpp

The companies you mention actually make LLMs, ornith-ai so far only make finetunes of existing LLMs, they re-trained existing LLMs with their own datasets or whatever, however the MTPs do not work correctly because they were made by Qwen for the oiriginal/official model, not for the onith-ai finetune(s).

Ironically the original Qwen3.6-35B MTP works because it's trained, while Ornith's is not. So your statement is wrong on both fact and reason.

Sign up or log in to comment