ComfyUI Qwen3-VL GGUF (8B & 4B)

GGUF quantizations and unquantized multimodal projectors (mmproj) extracted from Comfy-Org's Qwen3-VL repackaged weights.

These artifacts enable lightweight visual language reasoning across ComfyUI-GGUF, llama.cpp, llama-server, and Ollama on consumer hardware and CPU runtimes.


Model Overview

  • Original Weights Source: Comfy-Org/Qwen3-VL
  • Underlying Architecture: Alibaba Cloud Qwen3-VL (Qwen3VLForConditionalGeneration)
  • Projector Format: Standalone F16 GGUF (mmproj) preserving full vision patch fidelity
  • Quantization Engine: llama.cpp (b11209) k-quants

Available Files & Specifications

8B Series

File Name Precision File Size Description
mmproj-qwen3vl-8b-f16.gguf F16 1.08 GB Multimodal vision projector (required for image reasoning)
qwen3vl-8b-Q4_K_M.gguf Q4_K_M 4.68 GB Optimal balance of generation speed, memory, and accuracy
qwen3vl-8b-Q5_K_M.gguf Q5_K_M 5.45 GB Balanced higher-precision text backbone
qwen3vl-8b-Q6_K.gguf Q6_K 6.26 GB Near-lossless language reasoning and instruction following

4B Series

File Name Precision File Size Description
mmproj-qwen3vl-4b-f16.gguf F16 0.78 GB 4B Multimodal vision projector
qwen3vl-4b-Q4_K_M.gguf Q4_K_M 2.33 GB Ultra-lightweight edge quant for low-memory environments
qwen3vl-4b-Q5_K_M.gguf Q5_K_M 2.69 GB Precision 4B variant

Note on Projectors (mmproj): The vision projector maps image patches into the language model's latent embedding space. It is kept in unquantized F16 to prevent color shifts, bounding box degradation, or spatial grounding errors.


Quickstart & Usage

1. ComfyUI Setup (via ComfyUI-GGUF / ComfyUI-QwenVL)

  1. Place your desired language GGUF (e.g., qwen3vl-8b-Q4_K_M.gguf) into:
    ComfyUI/models/unet/
    # or
    ComfyUI/models/text_encoders/
    
  2. Place the matching projector (e.g., mmproj-qwen3vl-8b-f16.gguf) into:
    ComfyUI/models/clip_vision/
    
  3. Load the model using standard UnetLoaderGGUF or DualCLIPLoaderGGUF nodes in your workflow.

Acknowledgements

Downloads last month
2,564
GGUF
Model size
4B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Abiray/ComfyUI-Qwen3-VL-GGUF

Quantized
(1)
this model

Space using Abiray/ComfyUI-Qwen3-VL-GGUF 1