Instructions to use delebash/chatterbox-turbo-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use delebash/chatterbox-turbo-GGUF with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox Turbo β GGUF for audio.cpp, with voice cloning
8-bit and 16-bit GGUF conversions of ResembleAI/chatterbox-turbo
(revision 749d1c1a46eb10492095d68fbcf55691ccf137cd) for audio.cpp.
Resemble AI's distilled 350M Chatterbox: a GPT-2 medium T3, a GPT-2 BPE tokenizer with 19 inline tags, and a 2-step meanflow S3Gen decoder. English.
This is an unofficial conversion. The model, its weights and its licence are Resemble AI's. audio.cpp's own model repo has a Turbo package, repacked from a third-party GGUF without the encoders, so it speaks only its built-in voice. These files keep the voice encoder, the S3 speech tokenizer and the CAMPPlus speaker encoder, so the model clones a voice from a reference clip as well as speaking its built-in voice.
The files
| File | Size | SHA-256 |
|---|---|---|
chatterbox-turbo-q8_0.gguf |
880,908,100 bytes | c5d04a634c5534d8a5e53ed4644a5c40f73f250e5ba60847350c3661bf08b128 |
chatterbox-turbo-f16.gguf |
1,391,709,474 bytes | 35df71aef5814797ab0d33244cda863ff1897f5100f2bbfd635d31525c542af1 |
Standalone GGUFs: the audio.cpp model spec and the tokenizer files are embedded.
Which audio.cpp runs them
Voice cloning, and Nano's 12 attention heads, need JustVoice's copy of audio.cpp,
delebash/audio.cpp, from release v0.9.0-jv.3 on.
Usage
audiocpp_cli --task tts --family chatterbox_turbo --model chatterbox-turbo-q8_0.gguf --backend cuda \
--voice-ref speaker.wav --text "The harbour lights came on one by one." --out out.wav
Leave out --voice-ref for the built-in voice. The reference clip must be longer than 5
seconds. As upstream's tts_turbo.py does, it is loudness-normalised to β27 LUFS, the first
15 s give the T3's 375-token prompt and the first 10 s the decoder's.
How it was made
With convert_chatterbox_turbo.py, from Resemble's own files
(t3_turbo_v1.safetensors, s3gen_meanflow.safetensors, ve.safetensors, conds.pt and the tokenizer files),
and audiocpp_gguf built from the same commit:
python3 tools/community_models/chatterbox_turbo/convert_chatterbox_turbo.py \
--checkpoint <snapshot> --output chatterbox-turbo-q8_0.gguf --type q8_0
The 16-bit file is --type f16 (audio.cpp ships core Chatterbox at f16). No PyTorch is used:
the built-in voice is read from conds.pt directly. The three encoders are byte-identical to
core Chatterbox's own (ResembleAI/chatterbox's ve.safetensors and s3gen.safetensors), so
audio.cpp's core Chatterbox code computes a voice from them. The rest is renamed to what
audio.cpp's Turbo loaders read: GPT-2's Conv1D weights transposed, the vocoder's weight norm
folded, the head count stored as t3/hparams.num_heads.
What was checked
On a CPU build of that commit, one English line, seed 7, q8_0. The references were two
synthetic voices (Kokoro's af_heart, a US woman, and bm_george, a UK man; 9.5 s and 10.3 s).
"Similarity" is the cosine of Resemble's voice-encoder embedding against each reference: own /
other. Qwen3-ASR 1.7B read every render back.
| Render | Similarity | Median pitch | Read-back |
|---|---|---|---|
| Built-in voice | β | 220 Hz | exact |
| Clone of the US woman | 0.940 / 0.604 | 205 Hz (reference 202) | exact |
| Clone of the UK man | 0.923 / 0.593 | 146 Hz (reference 142) | exact |
Clones 0.935 and 0.918 at f16. A clip of 3 seconds is refused by name. Not checked: long-form output, listening
tests, real human voices, GPU speed.
Licence
MIT, the original's licence; the full text is in LICENSE. Changes from the original: the
weights were converted to GGUF β quantised to q8_0 or stored as f16 β with the tensors
renamed and the vocoder's weight norm folded as described above, and the tokenizer files and the
built-in voice embedded in the GGUF.
- Downloads last month
- 151
8-bit
16-bit
Model tree for delebash/chatterbox-turbo-GGUF
Base model
ResembleAI/chatterbox-turbo