GLM-5.3 GGUF

GGUF quantizations of zai-org/GLM-5.3-BF16, produced from BF16 GGUF weights with an importance matrix.

Quantizations

Quantization BPW Size Notes
Q4_K_M 4.83 454.85 GB Recommended quality/size balance; importance-matrix calibrated
IQ3_XXS 3.15 297.09 GB Compact 3-bit build; importance-matrix calibrated
IQ2_XS 2.42 227.69 GB Compact 2-bit build; importance-matrix calibrated

The shared indexer and the full NextN/draft block are kept at Q8_0 where the calibration matrix has no matching entries. This avoids unsafe low-bit fallback while retaining importance-matrix mixed quantization for the main model.

Usage

llama-cli -m Q4_K_M/GLM-5.3-Q4_K_M-00001-of-00033.gguf

Use a llama.cpp build with GLM-5.3/DeepSeek-V4 architecture support.

The original model is distributed under the GLM-5.3 license.

Downloads last month
136
GGUF
Model size
754B params
Architecture
glm-dsa
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0ppxnhximxr/GLM-5.3-GGUF

Quantized
(30)
this model