Chatterbox Turbo β€” GGUF for audio.cpp, with voice cloning

8-bit and 16-bit GGUF conversions of ResembleAI/chatterbox-turbo (revision 749d1c1a46eb10492095d68fbcf55691ccf137cd) for audio.cpp. Resemble AI's distilled 350M Chatterbox: a GPT-2 medium T3, a GPT-2 BPE tokenizer with 19 inline tags, and a 2-step meanflow S3Gen decoder. English.

This is an unofficial conversion. The model, its weights and its licence are Resemble AI's. audio.cpp's own model repo has a Turbo package, repacked from a third-party GGUF without the encoders, so it speaks only its built-in voice. These files keep the voice encoder, the S3 speech tokenizer and the CAMPPlus speaker encoder, so the model clones a voice from a reference clip as well as speaking its built-in voice.

The files

File Size SHA-256
chatterbox-turbo-q8_0.gguf 880,908,100 bytes c5d04a634c5534d8a5e53ed4644a5c40f73f250e5ba60847350c3661bf08b128
chatterbox-turbo-f16.gguf 1,391,709,474 bytes 35df71aef5814797ab0d33244cda863ff1897f5100f2bbfd635d31525c542af1

Standalone GGUFs: the audio.cpp model spec and the tokenizer files are embedded.

Which audio.cpp runs them

Voice cloning, and Nano's 12 attention heads, need JustVoice's copy of audio.cpp, delebash/audio.cpp, from release v0.9.0-jv.3 on.

Usage

audiocpp_cli --task tts --family chatterbox_turbo --model chatterbox-turbo-q8_0.gguf --backend cuda \
  --voice-ref speaker.wav --text "The harbour lights came on one by one." --out out.wav

Leave out --voice-ref for the built-in voice. The reference clip must be longer than 5 seconds. As upstream's tts_turbo.py does, it is loudness-normalised to βˆ’27 LUFS, the first 15 s give the T3's 375-token prompt and the first 10 s the decoder's.

How it was made

With convert_chatterbox_turbo.py, from Resemble's own files (t3_turbo_v1.safetensors, s3gen_meanflow.safetensors, ve.safetensors, conds.pt and the tokenizer files), and audiocpp_gguf built from the same commit:

python3 tools/community_models/chatterbox_turbo/convert_chatterbox_turbo.py \
  --checkpoint <snapshot> --output chatterbox-turbo-q8_0.gguf --type q8_0

The 16-bit file is --type f16 (audio.cpp ships core Chatterbox at f16). No PyTorch is used: the built-in voice is read from conds.pt directly. The three encoders are byte-identical to core Chatterbox's own (ResembleAI/chatterbox's ve.safetensors and s3gen.safetensors), so audio.cpp's core Chatterbox code computes a voice from them. The rest is renamed to what audio.cpp's Turbo loaders read: GPT-2's Conv1D weights transposed, the vocoder's weight norm folded, the head count stored as t3/hparams.num_heads.

What was checked

On a CPU build of that commit, one English line, seed 7, q8_0. The references were two synthetic voices (Kokoro's af_heart, a US woman, and bm_george, a UK man; 9.5 s and 10.3 s). "Similarity" is the cosine of Resemble's voice-encoder embedding against each reference: own / other. Qwen3-ASR 1.7B read every render back.

Render Similarity Median pitch Read-back
Built-in voice β€” 220 Hz exact
Clone of the US woman 0.940 / 0.604 205 Hz (reference 202) exact
Clone of the UK man 0.923 / 0.593 146 Hz (reference 142) exact

Clones 0.935 and 0.918 at f16. A clip of 3 seconds is refused by name. Not checked: long-form output, listening tests, real human voices, GPU speed.

Licence

MIT, the original's licence; the full text is in LICENSE. Changes from the original: the weights were converted to GGUF β€” quantised to q8_0 or stored as f16 β€” with the tensors renamed and the vocoder's weight norm folded as described above, and the tokenizer files and the built-in voice embedded in the GGUF.

Downloads last month
151
GGUF
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for delebash/chatterbox-turbo-GGUF

Quantized
(19)
this model