Neutrino-0.6B (speculative-decoding draft for Neutrino-8B)

This model is the speculative-decoding draft for Neutrino-8B. Paired with the 8B under greedy decoding, it makes the 8B faster while the output stays token-identical to plain 8B decoding. It is not an assistant: no chat tuning, no instruction following, no tool use. For a small model to talk to, use Neutrino-0.6B-Chat.

A 596M-parameter language model whose every transformer linear is stored in a ternary-family weight format eight times smaller than fp16, in a single 327,719,836-byte container (237,696,073 bytes to download), decoding at 225-236 tok/s on the CPU cores of a 16 GB Apple M5 with the same runtime binaries that execute Neutrino-8B. Draft and verifier use the same format and run in one process. The weights are trained to agree with Neutrino-8B's next-token choices.

  • Container: neutrino-0.6b_v4.bin, 28 layers, hidden 1024, vocabulary 151,936. 327,719,836 bytes, sha256 ba0e9258d866d5eb309352dd7c1b810c402eac289fb98e871c63b38a3bcd8168.
  • Compressed transport: neutrino-0.6b_v4.tv4z, 237,696,073 bytes, lossless, sha256 7ced43476235e6b7e2a0b5190a0fb3c59ae36f2f44f739574078b28dc50e6b8c.
  • 196 packed linears plus one int8 embedding table (tied: the input embedding and the output head share one table). The weights stay packed in memory and are decoded inside the matrix kernels.
  • Tokenizer shipped in this repo; vocabulary identical to the 8B's, which drafting requires.

Architecture

field value
Parameters 596,049,920 (440,401,920 packed projection + 155,582,464 int8 embedding + 65,536 fp32 norm)
Decoder layers 28
Hidden width 1,024
Feed-forward width 3,072, gated (SwiGLU)
Attention grouped-query 2:1, 16 query heads, 8 KV heads, head_dim 128
Rotary embedding full head width (rotary_dims 128), theta 1,000,000
Normalization RMSNorm, eps 1e-6, plus per-head Q/K RMSNorm in attention
Context length 40,960 tokens
Vocabulary 151,936
Embeddings tied, one int8 table serves the input embedding and the output head
KV cache 112 KiB/token fp16: 0.47 GB at 4k, 3.76 GB at 32k

Files

artifact bytes sha256
neutrino-0.6b_v4.bin (the container every runtime executes; the MLX pack reads this same file) 327,719,836 ba0e9258d866d5eb309352dd7c1b810c402eac289fb98e871c63b38a3bcd8168
neutrino-0.6b_v4.tv4z (lossless compressed transport, 72.5% of the container) 237,696,073 7ced43476235e6b7e2a0b5190a0fb3c59ae36f2f44f739574078b28dc50e6b8c
gguf/neutrino-0.6b-base-fv5.gguf (llama.cpp pack) see gguf/SHA256SUMS see gguf/SHA256SUMS
gguf/fv5.patch (the llama.cpp patch that builds the fork) - 81bf7be5eef18eb98ab4823596374c8dd6615151cf5615425973cb65942c0f73

MANIFEST.json lists the size and sha256 of every file in this repository.

Quickstart

Five ways to run it, all executing the same container.

1. pip engine (one command; pulls the container and the platform-matching native runtime):

pip install fermion-research
fermion generate --model fermionresearch/Neutrino-0.6B \
    --raw --max-new 64 "2, 3, 5, 7, 11,"  # base model: --raw skips the chat template; the prompt is positional

The first run downloads the ~0.3 GB container; later runs load from the local cache.

2. transformers (the reference path; registers as a native transformers model). Download the repo first: the loader resolves the container relative to a local path, so pass the downloaded directory to from_pretrained:

hf download fermionresearch/Neutrino-0.6B --local-dir Neutrino-0.6B \
    --exclude "gguf/*" --exclude "*.tv4z"   # skip the packs other runtimes use
import fermion  # registers the trtc_v4 model type
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Neutrino-0.6B")   # the downloaded directory
tokenizer = AutoTokenizer.from_pretrained("Neutrino-0.6B")
ids = tokenizer("2, 3, 5, 7, 11,", return_tensors="pt").input_ids
print(tokenizer.decode(model.generate(ids, max_new_tokens=64)[0]))

This is a base model: prompt it as a completion engine. Greedy decoding repeats itself on open-ended prompts, as 0.6B base models do; use sampling for anything user-facing.

import fermion must come first: it registers the trtc_v4 model type. Without it Transformers raises model type 'trtc_v4' ... not recognize this architecture and advises upgrading Transformers or installing it from source; neither helps, because the type is registered at import time by the fermion-research package, and trust_remote_code=True does not help because this repo carries no auto_map. Add the import.

3. Native binary (the fast path; this repo ships the fermion-run-* binaries under bin/; they run any Neutrino container, so the same binaries also run the 8B). From the downloaded repo directory:

(cd bin && shasum -a 256 -c fermion-run-macos-arm64.sha256)   # sidecar holds a bare filename
chmod +x bin/fermion-run-macos-arm64                          # hf download writes 0644
xattr -d com.apple.quarantine bin/fermion-run-macos-arm64 2>/dev/null || true  # macOS only
./bin/fermion-run-macos-arm64 neutrino-0.6b_v4.bin "1,791,1283" 128 9   # ids, n_new, threads

On Linux use bin/fermion-run-linux-x64 and skip the xattr line. The binaries' command line is documented in bin/README.md.

4. GGUF pack + our llama.cpp fork (also the draft file for llama.cpp's speculative decoding against the 8B pack). The fork is public at fermionresearch/llama.cpp, branch fermion-fv5 (= upstream ggml-org/llama.cpp @ d67c0b41 + the FV5 patch; gguf/fv5.patch in this repo is the same change as a mailbox patch):

hf download fermionresearch/Neutrino-0.6B gguf/neutrino-0.6b-base-fv5.gguf --local-dir .
git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp
git checkout fermion-fv5
cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF  # no FV5 Metal kernel yet: CPU path
cmake --build build -j --target llama-completion
./build/bin/llama-completion -m ../gguf/neutrino-0.6b-base-fv5.gguf \
    -p "2, 3, 5, 7, 11," -n 128 --temp 0 -no-cnv

5. MLX pack (Apple silicon, Metal kernels). Run it from the mlx/ folder of the downloaded repo, where the fermion_mlx package and its tokenizer files ship; requirements.txt installs mlx itself:

cd Neutrino-0.6B/mlx        # the directory downloaded in step 2
pip install -r requirements.txt
python -m fermion_mlx --model ../neutrino-0.6b_v4.bin --mode gen \
    --tokenizer . --prompt "2, 3, 5, 7, 11,"

Recommended settings, by use case

use case settings why
draft model for Neutrino-8B inherit the verifier; set nothing The draft's sampler cannot change what is emitted. Under greedy decoding the verifier's own choice decides acceptance. Tuning the draft changes throughput, never quality.
completion greedy, no penalty, no top-k (this repo's generation_config.json) A bare model.generate(), an lm-eval run and fermion generate all reproduce each other.
user-facing chat use Neutrino-0.6B-Chat This model has no chat tuning.

generation_config.json here is do_sample: false with no temperature, top-k, top-p or repetition penalty, so nothing silently overrides a harness. Top-k is not exposed on any runtime we ship; llama.cpp and ollama default to top-k 40, which is why the Modelfile in gguf/ sets it to 0.

Speed

Surface Machine Rate
native fermion-run, CPU only Apple M5, 9 threads 225-236 tok/s
MLX pack Apple M5, 16 GB 201 tok/s, 0.53 GiB peak memory
GGUF pack via fermion-fv5 CPU, 16 threads 55.4 tok/s
draft running inside the 8B's process H100 1,177 tok/s

As the 8B's draft

A 0.6B draft proposes k tokens; the 8B checks all k in one batched forward pass and keeps the accepted prefix plus one token of its own. A proposal is accepted only when it equals the 8B's own greedy choice, so the output is the plain greedy output: the speed changes, the text does not.

prompt class agreement with the 8B end-to-end speedup (H100)
counting 100% ×1.80
facts ~87% ×1.23
prose ~75% ×1.12
chat / explanation ~75% ×1.07
code ~78% ×1.06

Pair it with fermion chat --model fermionresearch/Neutrino-8B --draft fermionresearch/Neutrino-0.6B, the MLX pack's --mode spec (see mlx/README.md in the Neutrino-8B repo), the native binaries' --draft, or llama-speculative -md on our llama.cpp fork (see gguf/README.md in the Neutrino-8B repo).

License and attribution

Weights: Apache-2.0. Based on Qwen/Qwen3-0.6B (Apache-2.0, Alibaba Cloud); see LICENSE and NOTICE.

bin/ binaries and the compiled kernels in mlx/fermion_mlx/: prebuilt, free to use with these weights; no redistribution outside this repo; no reverse engineering.

Contact: contact@fermionresearch.com

Downloads last month
520
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FermionResearch/Neutrino-0.6B

Finetuned
Qwen/Qwen3-0.6B
Quantized
(476)
this model