OmniVoice β GGUF (ggml-quantised)
GGUF / ggml conversion of k2-fsa/OmniVoice
for use with CrispASR.
OmniVoice is a zero-shot multilingual TTS system with voice cloning and voice design. It is a Qwen3-0.6B AR backbone plus a HiggsAudioV2 audio tokenizer/codec (HuBERT semantic encoder + DAC acoustic encoder/decoder + RVQ).
Non-commercial. The upstream pre-trained model is CC-BY-NC (training data, e.g. Emilia), and the audio tokenizer is under the Boson Higgs Audio 2 Community License (Meta Llama 3 Community License lineage). See License & attribution before using or redistributing.
Files
| File | Quant | Size | Notes |
|---|---|---|---|
omnivoice-q8_0.gguf |
Q8_0 | 0.82 GB | Main model (AR backbone + audio heads) β recommended |
omnivoice-tokenizer-f32.gguf |
F32 | 0.81 GB | Audio tokenizer / codec β F32 for clone-encode fidelity |
Two files are needed: the main model (text β audio codes) and the tokenizer (audio codes β waveform, and reference audio β codes for cloning).
Which quant should I use?
The two halves have different sensitivity:
- Tokenizer β keep it wide (F16/F32). Voice cloning encodes the reference through an RVQ nearest-neighbour search; F16 rounding can flip near-boundary codebook picks, so a q8 tokenizer decodes fine but clones worse. This repo ships F32 (~0.81 GB) so the clone encode is the exact reference.
- Main model β q8_0 is near-lossless. The AR backbone tolerates Q8_0 with no audible loss and saves ~0.4 GB vs F16. Spend precision on the tokenizer, not here.
Recommended: omnivoice-q8_0.gguf + omnivoice-tokenizer-f32.gguf β
best quality per GB, and roughly the same footprint as an f16/f16 pairing.
Other builds can be produced from the upstream F32 checkpoint with
models/convert-omnivoice-to-gguf.py --outtype f32|f16 and
models/convert-omnivoice-tokenizer-to-gguf.py --outtype f32|f16
(then crispasr-quantize in.gguf out-q8_0.gguf q8_0).
Quick start
# 1. Build CrispASR (CUDA build shown; CPU also works)
git clone https://github.com/CrispStrobe/CrispASR
cd CrispASR
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j --target crispasr-cli
# 2. Pull both files
hf download Bonenk/OmniVoice-GGUF omnivoice-q8_0.gguf --local-dir models
hf download Bonenk/OmniVoice-GGUF omnivoice-tokenizer-f32.gguf --local-dir models
CLI β synthesise a line
./build/bin/crispasr --backend omnivoice \
-m models/omnivoice-q8_0.gguf \
--codec-model models/omnivoice-tokenizer-f32.gguf \
--tts "Hello from OmniVoice." \
--tts-output out.wav
Clone a voice from a reference clip (consent required β --i-have-rights
attests you have the speaker's permission or it is your own voice):
./build/bin/crispasr --backend omnivoice \
-m models/omnivoice-q8_0.gguf \
--codec-model models/omnivoice-tokenizer-f32.gguf \
--voice reference.wav --i-have-rights \
--tts "This is a cloned voice." --tts-output clone.wav
Add --tts-steps N (default 32) to trade speed for quality β it is the dominant
knob (stage0 cost β N Γ 2 forwards). --tts-duration SECONDS forces an exact
output length.
Server β OpenAI-compatible /v1/audio/speech
./build/bin/crispasr --server --backend omnivoice \
-m models/omnivoice-q8_0.gguf \
--codec-model models/omnivoice-tokenizer-f32.gguf \
--voice-dir voices --port 8000 \
--no-punctuation --accept-marking-responsibility
curl -s http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Hello from OmniVoice.","voice":"my_voice",
"consent_attestation":"I have the speaker'"'"'s consent"}' \
--output out.wav
voice is a name in --voice-dir (uploaded once via POST /v1/voices) or a
preset; a .wav clone requires consent_attestation.
Useful environment variables
| Variable | Default | Effect |
|---|---|---|
CRISPASR_OMNIVOICE_NUM_STEPS |
32 | masked-iterative steps (quality β speed) |
CRISPASR_OMNIVOICE_GUIDANCE |
2.0 | guidance scale |
CRISPASR_OMNIVOICE_CLASS_TEMP |
0 | 0 = greedy/deterministic |
CRISPASR_TTS_REF_CACHE_DIR |
temp | content-addressed reference-code cache |
Restricted weights may require accepting the upstream licence before download:
--accept-license cc-by-nc (or all).
Conversion
python models/convert-omnivoice-to-gguf.py \
--input k2-fsa/OmniVoice --output omnivoice-f32.gguf --outtype f32
python models/convert-omnivoice-tokenizer-to-gguf.py \
--input k2-fsa/OmniVoice/audio_tokenizer \
--output omnivoice-tokenizer-f32.gguf --outtype f32
License & attribution
This is a format conversion of upstream weights; it grants no additional rights. The upstream terms apply:
- Code: Apache-2.0.
- Pre-trained model (
k2-fsa/OmniVoice): CC-BY-NC β non-commercial use only, due to training-data constraints (e.g. Emilia). Attribution required. - Audio tokenizer (HiggsAudioV2): Boson Higgs Audio 2 Community License
Agreement, built on the Meta Llama 3 Community License. Redistribution is
permitted subject to its conditions, including including the licence and the
"Built with Meta Llama 3" attribution notice. See
LICENSE-tokenizer.txt.
Voice cloning / impersonation: users must have the speaker's consent and must comply with local law. Unauthorized voice cloning, impersonation, fraud or scams are prohibited. CrispASR gates cloning behind an explicit consent attestation and can attach an audible AI-disclosure to synthetic speech.
Credits
- Upstream model:
k2-fsa/OmniVoiceΒ· code: k2-fsa/OmniVoice - GGUF conversion & runtime: CrispStrobe/CrispASR
- Downloads last month
- 101
8-bit
32-bit