πŸš€ Nemotron-3-Embed-8B (GGUF Quantizations)

Model Architecture Context Length RTEB Rank License

Welcome to the GGUF repository for NVIDIA's Nemotron-3-Embed-8B-BF16.

This model is a state-of-the-art, 8-billion parameter multilingual text embedding model optimized for Retrieval-Augmented Generation (RAG), semantic search, and cross-lingual retrieval workflows. It achieved #1 on the multilingual RTEB leaderboard (as of July 2026).

By converting the model to GGUF (GPT-Generated Unified Format), you can run enterprise-grade retrieval locally on consumer hardware (CPUs and GPUs) using tools like llama.cpp, Ollama, or LM Studio.


πŸ“‚ Repository Files & Quantization Options

Below is the directory structure of the available .gguf files in this repository. Choose the quantization level that best fits your VRAM/RAM constraints!

πŸ“ nemotron-3-embed-8b-gguf/
β”‚
β”œβ”€β”€ πŸ“„ README.md
β”œβ”€β”€ πŸ“„ config.json
β”‚
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q8_0.gguf    (8.46 GB) 🟒 Near zero quality loss
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q6_K.gguf    (6.53 GB) 🟒 Extremely low quality loss
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q5_K_M.gguf  (5.69 GB) 🟑 Very low quality loss
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q5_K_S.gguf  (5.55 GB) 🟑 Very low quality loss
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q4_K_M.gguf  (4.90 GB) ⭐ RECOMMENDED - Great balance
β”œβ”€β”€ πŸ“¦ Nemotron-3-Embed-8B-Q4_K_S.gguf  (4.65 GB) 🟠 Moderate quality loss
└── πŸ“¦ Nemotron-3-Embed-8B-Q3_K_M.gguf  (4.01 GB) πŸ”΄ High quality loss (Memory constrained only)
Downloads last month
169,078
GGUF
Model size
8B params
Architecture
mistral3
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Abiray/Nemotron-3-Embed-8B-GGUF

Space using Abiray/Nemotron-3-Embed-8B-GGUF 1