Instructions to use FermionResearch/Neutrino-0.6B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FermionResearch/Neutrino-0.6B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FermionResearch/Neutrino-0.6B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("FermionResearch/Neutrino-0.6B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FermionResearch/Neutrino-0.6B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FermionResearch/Neutrino-0.6B # Run inference directly in the terminal: llama cli -hf FermionResearch/Neutrino-0.6B
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FermionResearch/Neutrino-0.6B # Run inference directly in the terminal: llama cli -hf FermionResearch/Neutrino-0.6B
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FermionResearch/Neutrino-0.6B # Run inference directly in the terminal: ./llama-cli -hf FermionResearch/Neutrino-0.6B
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FermionResearch/Neutrino-0.6B # Run inference directly in the terminal: ./build/bin/llama-cli -hf FermionResearch/Neutrino-0.6B
Use Docker
docker model run hf.co/FermionResearch/Neutrino-0.6B
- LM Studio
- Jan
- vLLM
How to use FermionResearch/Neutrino-0.6B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FermionResearch/Neutrino-0.6B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FermionResearch/Neutrino-0.6B
- SGLang
How to use FermionResearch/Neutrino-0.6B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FermionResearch/Neutrino-0.6B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FermionResearch/Neutrino-0.6B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-0.6B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use FermionResearch/Neutrino-0.6B with Ollama:
ollama run hf.co/FermionResearch/Neutrino-0.6B
- Unsloth Desktop
- Pi
How to use FermionResearch/Neutrino-0.6B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-0.6B
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FermionResearch/Neutrino-0.6B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FermionResearch/Neutrino-0.6B with Docker Model Runner:
docker model run hf.co/FermionResearch/Neutrino-0.6B
- Lemonade
How to use FermionResearch/Neutrino-0.6B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FermionResearch/Neutrino-0.6B
Run and chat with the model
lemonade run user.Neutrino-0.6B-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use FermionResearch/Neutrino-0.6B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-0.6B
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FermionResearch/Neutrino-0.6B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FermionResearch/Neutrino-0.6B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-0.6B
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FermionResearch/Neutrino-0.6B" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Neutrino-0.6B (speculative-decoding draft for Neutrino-8B)
This model is the speculative-decoding draft for Neutrino-8B. Paired with the 8B under greedy decoding, it makes the 8B faster while the output stays token-identical to plain 8B decoding. It is not an assistant: no chat tuning, no instruction following, no tool use. For a small model to talk to, use Neutrino-0.6B-Chat.
A 596M-parameter language model whose every transformer linear is stored in a ternary-family weight format eight times smaller than fp16, in a single 327,719,836-byte container (237,696,073 bytes to download), decoding at 225-236 tok/s on the CPU cores of a 16 GB Apple M5 with the same runtime binaries that execute Neutrino-8B. Draft and verifier use the same format and run in one process. The weights are trained to agree with Neutrino-8B's next-token choices.
- Container:
neutrino-0.6b_v4.bin, 28 layers, hidden 1024, vocabulary 151,936. 327,719,836 bytes, sha256ba0e9258d866d5eb309352dd7c1b810c402eac289fb98e871c63b38a3bcd8168. - Compressed transport:
neutrino-0.6b_v4.tv4z, 237,696,073 bytes, lossless, sha2567ced43476235e6b7e2a0b5190a0fb3c59ae36f2f44f739574078b28dc50e6b8c. - 196 packed linears plus one int8 embedding table (tied: the input embedding and the output head share one table). The weights stay packed in memory and are decoded inside the matrix kernels.
- Tokenizer shipped in this repo; vocabulary identical to the 8B's, which drafting requires.
Architecture
| field | value |
|---|---|
| Parameters | 596,049,920 (440,401,920 packed projection + 155,582,464 int8 embedding + 65,536 fp32 norm) |
| Decoder layers | 28 |
| Hidden width | 1,024 |
| Feed-forward width | 3,072, gated (SwiGLU) |
| Attention | grouped-query 2:1, 16 query heads, 8 KV heads, head_dim 128 |
| Rotary embedding | full head width (rotary_dims 128), theta 1,000,000 |
| Normalization | RMSNorm, eps 1e-6, plus per-head Q/K RMSNorm in attention |
| Context length | 40,960 tokens |
| Vocabulary | 151,936 |
| Embeddings | tied, one int8 table serves the input embedding and the output head |
| KV cache | 112 KiB/token fp16: 0.47 GB at 4k, 3.76 GB at 32k |
Files
| artifact | bytes | sha256 |
|---|---|---|
neutrino-0.6b_v4.bin (the container every runtime executes; the MLX pack reads this same file) |
327,719,836 | ba0e9258d866d5eb309352dd7c1b810c402eac289fb98e871c63b38a3bcd8168 |
neutrino-0.6b_v4.tv4z (lossless compressed transport, 72.5% of the container) |
237,696,073 | 7ced43476235e6b7e2a0b5190a0fb3c59ae36f2f44f739574078b28dc50e6b8c |
gguf/neutrino-0.6b-base-fv5.gguf (llama.cpp pack) |
see gguf/SHA256SUMS |
see gguf/SHA256SUMS |
gguf/fv5.patch (the llama.cpp patch that builds the fork) |
- | 81bf7be5eef18eb98ab4823596374c8dd6615151cf5615425973cb65942c0f73 |
MANIFEST.json lists the size and sha256 of every file in this repository.
Quickstart
Five ways to run it, all executing the same container.
1. pip engine (one command; pulls the container and the platform-matching native runtime):
pip install fermion-research
fermion generate --model fermionresearch/Neutrino-0.6B \
--raw --max-new 64 "2, 3, 5, 7, 11," # base model: --raw skips the chat template; the prompt is positional
The first run downloads the ~0.3 GB container; later runs load from the local cache.
2. transformers (the reference path; registers as a native
transformers model). Download the repo first: the loader resolves the
container relative to a local path, so pass the downloaded directory to
from_pretrained:
hf download fermionresearch/Neutrino-0.6B --local-dir Neutrino-0.6B \
--exclude "gguf/*" --exclude "*.tv4z" # skip the packs other runtimes use
import fermion # registers the trtc_v4 model type
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Neutrino-0.6B") # the downloaded directory
tokenizer = AutoTokenizer.from_pretrained("Neutrino-0.6B")
ids = tokenizer("2, 3, 5, 7, 11,", return_tensors="pt").input_ids
print(tokenizer.decode(model.generate(ids, max_new_tokens=64)[0]))
This is a base model: prompt it as a completion engine. Greedy decoding repeats itself on open-ended prompts, as 0.6B base models do; use sampling for anything user-facing.
import fermion must come first: it registers the trtc_v4 model type.
Without it Transformers raises model type 'trtc_v4' ... not recognize this architecture and advises upgrading Transformers or installing it from
source; neither helps, because the type is registered at import time by the
fermion-research package, and trust_remote_code=True does not help
because this repo carries no auto_map. Add the import.
3. Native binary (the fast path; this repo ships the fermion-run-*
binaries under bin/; they run any Neutrino container, so the same
binaries also run the 8B). From the downloaded repo directory:
(cd bin && shasum -a 256 -c fermion-run-macos-arm64.sha256) # sidecar holds a bare filename
chmod +x bin/fermion-run-macos-arm64 # hf download writes 0644
xattr -d com.apple.quarantine bin/fermion-run-macos-arm64 2>/dev/null || true # macOS only
./bin/fermion-run-macos-arm64 neutrino-0.6b_v4.bin "1,791,1283" 128 9 # ids, n_new, threads
On Linux use bin/fermion-run-linux-x64 and skip the xattr line. The
binaries' command line is documented in bin/README.md.
4. GGUF pack + our llama.cpp fork (also the draft file for llama.cpp's
speculative decoding against the 8B pack). The fork is public at
fermionresearch/llama.cpp,
branch fermion-fv5 (= upstream ggml-org/llama.cpp @ d67c0b41 + the
FV5 patch; gguf/fv5.patch in this repo is the same change as a mailbox
patch):
hf download fermionresearch/Neutrino-0.6B gguf/neutrino-0.6b-base-fv5.gguf --local-dir .
git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp
git checkout fermion-fv5
cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF # no FV5 Metal kernel yet: CPU path
cmake --build build -j --target llama-completion
./build/bin/llama-completion -m ../gguf/neutrino-0.6b-base-fv5.gguf \
-p "2, 3, 5, 7, 11," -n 128 --temp 0 -no-cnv
5. MLX pack (Apple silicon, Metal kernels). Run it from the mlx/
folder of the downloaded repo, where the fermion_mlx package and its
tokenizer files ship; requirements.txt installs mlx itself:
cd Neutrino-0.6B/mlx # the directory downloaded in step 2
pip install -r requirements.txt
python -m fermion_mlx --model ../neutrino-0.6b_v4.bin --mode gen \
--tokenizer . --prompt "2, 3, 5, 7, 11,"
Recommended settings, by use case
| use case | settings | why |
|---|---|---|
| draft model for Neutrino-8B | inherit the verifier; set nothing | The draft's sampler cannot change what is emitted. Under greedy decoding the verifier's own choice decides acceptance. Tuning the draft changes throughput, never quality. |
| completion | greedy, no penalty, no top-k (this repo's generation_config.json) |
A bare model.generate(), an lm-eval run and fermion generate all reproduce each other. |
| user-facing chat | use Neutrino-0.6B-Chat | This model has no chat tuning. |
generation_config.json here is do_sample: false with no temperature,
top-k, top-p or repetition penalty, so nothing silently overrides a
harness. Top-k is not exposed on any runtime we ship; llama.cpp and ollama
default to top-k 40, which is why the Modelfile in gguf/ sets it to 0.
Speed
| Surface | Machine | Rate |
|---|---|---|
native fermion-run, CPU only |
Apple M5, 9 threads | 225-236 tok/s |
| MLX pack | Apple M5, 16 GB | 201 tok/s, 0.53 GiB peak memory |
GGUF pack via fermion-fv5 |
CPU, 16 threads | 55.4 tok/s |
| draft running inside the 8B's process | H100 | 1,177 tok/s |
As the 8B's draft
A 0.6B draft proposes k tokens; the 8B checks all k in one batched forward pass and keeps the accepted prefix plus one token of its own. A proposal is accepted only when it equals the 8B's own greedy choice, so the output is the plain greedy output: the speed changes, the text does not.
| prompt class | agreement with the 8B | end-to-end speedup (H100) |
|---|---|---|
| counting | 100% | ×1.80 |
| facts | ~87% | ×1.23 |
| prose | ~75% | ×1.12 |
| chat / explanation | ~75% | ×1.07 |
| code | ~78% | ×1.06 |
Pair it with fermion chat --model fermionresearch/Neutrino-8B --draft fermionresearch/Neutrino-0.6B, the MLX pack's --mode spec (see
mlx/README.md in the Neutrino-8B repo), the native binaries' --draft,
or llama-speculative -md on our llama.cpp fork (see gguf/README.md in
the Neutrino-8B repo).
License and attribution
Weights: Apache-2.0. Based on Qwen/Qwen3-0.6B (Apache-2.0, Alibaba Cloud);
see LICENSE and NOTICE.
bin/ binaries and the compiled kernels in mlx/fermion_mlx/: prebuilt,
free to use with these weights; no redistribution outside this repo; no
reverse engineering.
Contact: contact@fermionresearch.com
- Downloads last month
- 520
We're not able to determine the quantization variants.