OmniVoice β€” GGUF (ggml-quantised)

GGUF / ggml conversion of k2-fsa/OmniVoice for use with CrispASR.

OmniVoice is a zero-shot multilingual TTS system with voice cloning and voice design. It is a Qwen3-0.6B AR backbone plus a HiggsAudioV2 audio tokenizer/codec (HuBERT semantic encoder + DAC acoustic encoder/decoder + RVQ).

Non-commercial. The upstream pre-trained model is CC-BY-NC (training data, e.g. Emilia), and the audio tokenizer is under the Boson Higgs Audio 2 Community License (Meta Llama 3 Community License lineage). See License & attribution before using or redistributing.

Files

File Quant Size Notes
omnivoice-q8_0.gguf Q8_0 0.82 GB Main model (AR backbone + audio heads) β€” recommended
omnivoice-tokenizer-f32.gguf F32 0.81 GB Audio tokenizer / codec β€” F32 for clone-encode fidelity

Two files are needed: the main model (text β†’ audio codes) and the tokenizer (audio codes β†’ waveform, and reference audio β†’ codes for cloning).

Which quant should I use?

The two halves have different sensitivity:

  • Tokenizer β€” keep it wide (F16/F32). Voice cloning encodes the reference through an RVQ nearest-neighbour search; F16 rounding can flip near-boundary codebook picks, so a q8 tokenizer decodes fine but clones worse. This repo ships F32 (~0.81 GB) so the clone encode is the exact reference.
  • Main model β€” q8_0 is near-lossless. The AR backbone tolerates Q8_0 with no audible loss and saves ~0.4 GB vs F16. Spend precision on the tokenizer, not here.

Recommended: omnivoice-q8_0.gguf + omnivoice-tokenizer-f32.gguf β€” best quality per GB, and roughly the same footprint as an f16/f16 pairing.

Other builds can be produced from the upstream F32 checkpoint with models/convert-omnivoice-to-gguf.py --outtype f32|f16 and models/convert-omnivoice-tokenizer-to-gguf.py --outtype f32|f16 (then crispasr-quantize in.gguf out-q8_0.gguf q8_0).

Quick start

# 1. Build CrispASR (CUDA build shown; CPU also works)
git clone https://github.com/CrispStrobe/CrispASR
cd CrispASR
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j --target crispasr-cli

# 2. Pull both files
hf download Bonenk/OmniVoice-GGUF omnivoice-q8_0.gguf --local-dir models
hf download Bonenk/OmniVoice-GGUF omnivoice-tokenizer-f32.gguf --local-dir models

CLI β€” synthesise a line

./build/bin/crispasr --backend omnivoice \
    -m  models/omnivoice-q8_0.gguf \
    --codec-model models/omnivoice-tokenizer-f32.gguf \
    --tts "Hello from OmniVoice." \
    --tts-output out.wav

Clone a voice from a reference clip (consent required β€” --i-have-rights attests you have the speaker's permission or it is your own voice):

./build/bin/crispasr --backend omnivoice \
    -m  models/omnivoice-q8_0.gguf \
    --codec-model models/omnivoice-tokenizer-f32.gguf \
    --voice reference.wav --i-have-rights \
    --tts "This is a cloned voice." --tts-output clone.wav

Add --tts-steps N (default 32) to trade speed for quality β€” it is the dominant knob (stage0 cost β‰ˆ N Γ— 2 forwards). --tts-duration SECONDS forces an exact output length.

Server β€” OpenAI-compatible /v1/audio/speech

./build/bin/crispasr --server --backend omnivoice \
    -m  models/omnivoice-q8_0.gguf \
    --codec-model models/omnivoice-tokenizer-f32.gguf \
    --voice-dir voices --port 8000 \
    --no-punctuation --accept-marking-responsibility
curl -s http://127.0.0.1:8000/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{"input":"Hello from OmniVoice.","voice":"my_voice",
       "consent_attestation":"I have the speaker'"'"'s consent"}' \
  --output out.wav

voice is a name in --voice-dir (uploaded once via POST /v1/voices) or a preset; a .wav clone requires consent_attestation.

Useful environment variables

Variable Default Effect
CRISPASR_OMNIVOICE_NUM_STEPS 32 masked-iterative steps (quality ↔ speed)
CRISPASR_OMNIVOICE_GUIDANCE 2.0 guidance scale
CRISPASR_OMNIVOICE_CLASS_TEMP 0 0 = greedy/deterministic
CRISPASR_TTS_REF_CACHE_DIR temp content-addressed reference-code cache

Restricted weights may require accepting the upstream licence before download: --accept-license cc-by-nc (or all).

Conversion

python models/convert-omnivoice-to-gguf.py \
    --input k2-fsa/OmniVoice --output omnivoice-f32.gguf --outtype f32
python models/convert-omnivoice-tokenizer-to-gguf.py \
    --input k2-fsa/OmniVoice/audio_tokenizer \
    --output omnivoice-tokenizer-f32.gguf --outtype f32

License & attribution

This is a format conversion of upstream weights; it grants no additional rights. The upstream terms apply:

  • Code: Apache-2.0.
  • Pre-trained model (k2-fsa/OmniVoice): CC-BY-NC β€” non-commercial use only, due to training-data constraints (e.g. Emilia). Attribution required.
  • Audio tokenizer (HiggsAudioV2): Boson Higgs Audio 2 Community License Agreement, built on the Meta Llama 3 Community License. Redistribution is permitted subject to its conditions, including including the licence and the "Built with Meta Llama 3" attribution notice. See LICENSE-tokenizer.txt.

Voice cloning / impersonation: users must have the speaker's consent and must comply with local law. Unauthorized voice cloning, impersonation, fraud or scams are prohibited. CrispASR gates cloning behind an explicit consent attestation and can attach an audible AI-disclosure to synthetic speech.

Credits

Downloads last month
101
GGUF
Model size
0.6B params
Architecture
omnivoice
Hardware compatibility
Log In to add your hardware

8-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Bonenk/OmniVoice-GGUF

Finetuned
Qwen/Qwen3-0.6B
Finetuned
k2-fsa/OmniVoice
Quantized
(40)
this model