Due to a mistake on my side, faulty files were in effect a few days after posting. If you downloaded any of these before October 9th, 2026, please re-download your GGUF file that has been corrected.

Qwen3.8 2B Distill

A distilled model with a base of Qwen3.5 2B fine-tuned with traces of a Qwen3.8 2.4T A95B teacher. Excels in benchmarks when compared to models of similar size. You can find more information about it here.

Standard quants

Quant Quant Size
FP16 3.9GB, highest accuracy
BF16 same size as FP16 with worse support for older graphics devices
Q8_0 2.08GB, no noticeable difference when compared to FP16/BF16
Q6_K 1.61GB, no noticeable difference when compared to Q8_0
Q5_K_M 1.45GB
Q5_K_S 1.42GB
Q4_K_M 1.31GB, the sweet spot between quality and speed
Q4_K_S 1.21GB
Q3_K_L 1.2GB
Q3_K_M 1.13GB
Q3_K_S 1.05GB
Q2_K 990MB, lowest accuracy on the list
Downloads last month
2,894
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for piskle/Qwen3.8-2B-Distill-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(22)
this model

Space using piskle/Qwen3.8-2B-Distill-GGUF 1