Instructions to use Abiray/Nemotron-3-Embed-8B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Abiray/Nemotron-3-Embed-8B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Abiray/Nemotron-3-Embed-8B-GGUF with Ollama:
ollama run hf.co/Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Abiray/Nemotron-3-Embed-8B-GGUF with Docker Model Runner:
docker model run hf.co/Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
- Lemonade
How to use Abiray/Nemotron-3-Embed-8B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Abiray/Nemotron-3-Embed-8B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Nemotron-3-Embed-8B-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
π Nemotron-3-Embed-8B (GGUF Quantizations)
Welcome to the GGUF repository for NVIDIA's Nemotron-3-Embed-8B-BF16.
This model is a state-of-the-art, 8-billion parameter multilingual text embedding model optimized for Retrieval-Augmented Generation (RAG), semantic search, and cross-lingual retrieval workflows. It achieved #1 on the multilingual RTEB leaderboard (as of July 2026).
By converting the model to GGUF (GPT-Generated Unified Format), you can run enterprise-grade retrieval locally on consumer hardware (CPUs and GPUs) using tools like llama.cpp, Ollama, or LM Studio.
π Repository Files & Quantization Options
Below is the directory structure of the available .gguf files in this repository. Choose the quantization level that best fits your VRAM/RAM constraints!
π nemotron-3-embed-8b-gguf/
β
βββ π README.md
βββ π config.json
β
βββ π¦ Nemotron-3-Embed-8B-Q8_0.gguf (8.46 GB) π’ Near zero quality loss
βββ π¦ Nemotron-3-Embed-8B-Q6_K.gguf (6.53 GB) π’ Extremely low quality loss
βββ π¦ Nemotron-3-Embed-8B-Q5_K_M.gguf (5.69 GB) π‘ Very low quality loss
βββ π¦ Nemotron-3-Embed-8B-Q5_K_S.gguf (5.55 GB) π‘ Very low quality loss
βββ π¦ Nemotron-3-Embed-8B-Q4_K_M.gguf (4.90 GB) β RECOMMENDED - Great balance
βββ π¦ Nemotron-3-Embed-8B-Q4_K_S.gguf (4.65 GB) π Moderate quality loss
βββ π¦ Nemotron-3-Embed-8B-Q3_K_M.gguf (4.01 GB) π΄ High quality loss (Memory constrained only)
- Downloads last month
- 169,078
3-bit
4-bit
5-bit
6-bit
8-bit
Model tree for Abiray/Nemotron-3-Embed-8B-GGUF
Base model
mistralai/Ministral-3-8B-Base-2512