CoCo-Decision-4B-GGUF

CoCo-Decision-4B is a specialized, 4.21B-parameter text-only decision model developed by Corners AI by fine-tuning Qwen3.5-4B to output probability distributions across multiple-choice options in a single forward pass without autoregressive text generation. Released under a CC BY-NC 4.0 license, it is engineered for high-throughput, low-latency pipeline tasks such as intent routing, triage, agent guardrails, and LLM-as-a-judge evaluation—achieving a median response latency of 26.1 ms on an NVIDIA RTX 5090. As of October 2026, it ranked 3rd out of 65 sub-5B models on the public Decision Index 0.3 benchmark (scoring 47.80, with particular strength in tools, automation, and language understanding), and it can be deployed via standard Hugging Face transformers or served through a /v1/systemone endpoint using oh-my-jev (omj).

Model Files

File Name Quant Type File Size File Link Description
CoCo-Decision-4B.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
CoCo-Decision-4B.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
CoCo-Decision-4B.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
CoCo-Decision-4B.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
CoCo-Decision-4B.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
CoCo-Decision-4B.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
CoCo-Decision-4B.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
CoCo-Decision-4B.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Releases / v0.6.0 — https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0

Downloads last month
479
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/CoCo-Decision-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(2)
this model

Collections including prithivMLmods/CoCo-Decision-4B-GGUF