Quipu 0.8B — GGUF

GGUF quantized versions of Quipuai/quipu-0.8b, ready to run with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

Quipu 0.8B is a LoRA fine-tune of Qwen3.5-0.8B, tuned for step-by-step reasoning, honest tool-calling behavior, consistent identity, and lightweight coding help.

Available quantizations

Quant Size Notes
Q2_K 422 MB Smallest, noticeable quality loss
Q3_K_S / Q3_K_M / Q3_K_L 435–491 MB Small, usable for simple tasks
Q4_0 / Q4_1 501–533 MB Legacy 4-bit, works everywhere
Q4_K_S / Q4_K_M 505–529 MB Recommended default — best size/quality balance
Q5_0 / Q5_1 564–595 MB Better quality, still compact
Q5_K_S / Q5_K_M 564–578 MB Good quality, low loss
Q6_K — Near-lossless
Q8_0 — Best quality, closest to full precision
F16 — Full precision, reference version

At 0.8B parameters the model is already tiny, so Q4_K_M or Q8_0 are the sweet spots — Q4_K_M if you want the smallest footprint with minimal loss, Q8_0 if you want the closest to full quality and don't mind a slightly bigger file.

Running with Ollama

ollama run hf.co/Quipuai/quipu-0.8b-GGUF:Q4_K_M

Swap Q4_K_M for any quant tag from the table above.

Running with llama.cpp

./llama-cli -m quipu-0.8b-Q4_K_M.gguf -p "Who are you?"

Full precision / training

For the unquantized safetensors model (for fine-tuning or merging), see Quipuai/quipu-0.8b.

Downloads last month
198
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Quipuai/quipu-0.8b-GGUF

Adapter
(2)
this model