LFM2.5-Audio, French anchoring: experiment checkpoints (plan v3, October 2026)
Intermediate checkpoints of two fine-tuning experiments of LiquidAI/LFM2.5-Audio-1.5B towards a French voice assistant (French in the SIWIS voice, English in the base model's "US female" voice, while preserving the text abilities of the language model). They are experiment artefacts, not a release: they are published so that the training machine's disk can be freed without losing the weights, and so that any of them can serve as the starting point of a new run.
Layout
runs/<run>/checkpoint_N/model.safetensors weights at an intermediate step (+ config.json of the base model)
runs/<run>/final/model.safetensors weights at the last step
runs/<run>/params.json full training configuration written by the training tool
runs/<run>/eval/... metrics (summary.json, per-sentence CSV), no audio
Only the model weights are kept. Optimizer, scheduler and RNG states were not uploaded, so a run cannot be
resumed exactly; a new run can start from any of these weights. The tokenizer, processor and audio detokenizer are
those of the base model (LiquidAI/LFM2.5-Audio-1.5B@c362a0625dfe45aa588dce5f0ada28a7e5707628): download one
folder, then either pass it to s2s-finetune --model-id <folder> in the training repository, or copy the base
model's other files next to model.safetensors and load the folder with LFM2AudioModel.from_pretrained.
Recommended decoding (decision D42 of the project): sample the audio with temperature 1.0 and top-k 4, in sequential and interleaved generation alike; keep the text greedy.
Common data
Both runs build on the corpus anchor_v5 (207,033 training examples, context 1,024 tokens). Every example
supervises the assistant turn only.
- ASR on real multi-speaker speech: French from Emilia-YODAS, Multilingual LibriSpeech, Common Voice,
VoxPopuli and CEFC-Orféo (31,056 clips); English from LibriTTS-R (21,211 clips). Prompts
Perform ASR. The audio is in French./Perform ASR. The audio is in English. - TTS and interleaved TTS (6 text tokens / 12 audio frames) on the assistant voice only: SIWIS studio
recordings (9,750 clips, 10.7 h;
Rcarvalo/siwis-fr-24k) and the assistant replies of synthetic dialogues re-synthesised in the target voice (French: 21,267 clips / 50.6 h with a Qwen3-TTS fine-tuned on SIWIS; English: 21,216 clips / 48.4 h with the base model's US female voice). PromptsPerform TTS. Use the French female voice.,Perform TTS with interleaved text and audio. Use the French female voice.(and... the US female voice.in English). - Text only (about half of the examples): 39,198 French and 14,727 English question/answer exchanges
(
Rcarvalo/text-replay-fr-en), seen twice, promptsTu es un assistant utile. Réponds en français, de façon claire et exacte./You are a helpful assistant. Answer in English, clearly and accurately.
Probe sentences and the validation/test dialogues of the next phase are excluded; splits are grouped by recording or dialogue.
Run stage_a: "language" step, whole model at full speed
| Start | base model LiquidAI/LFM2.5-Audio-1.5B@c362a06 |
| Corpus | anchor_v5 + 50,000 real French Emilia-YODAS clips used as TTS / interleaved TTS targets under the generic prompt Perform TTS. Use a French voice. (many speakers; SIWIS keeps its named prompt) |
| Learning rate | 1e-4 for every parameter (lm_lr_scale 1.0), 150 warm-up steps, cosine decay |
| Batch | 4 x 16 accumulation = 64 examples, bf16, AdamW, gradient clipping 1.0 |
| Steps | 2,000; checkpoints every 500 steps (checkpoint_0 = step 500 ... checkpoint_3 = step 2,000 = final) |
Goal: learn French phonetics quickly from real multi-speaker speech before a second, slow step for the voice.
| Checkpoint (step) | INCLUDE-fr (419 q.) | French TTS WER |
|---|---|---|
| base model | 42.2 % | 0.62 |
checkpoint_0 (500) |
39.1 % | 0.319 |
checkpoint_1 (1,000) |
33.7 % | 0.262 |
checkpoint_2 (1,500) |
37.7 % | 0.190 |
checkpoint_3 (2,000) |
37.7 % | 0.217 |
TTS: 50 fixed French probe sentences never seen in training, D42 sampling, seed 0, generic prompt
Perform TTS. Use a French voice., transcribed by Whisper-small. INCLUDE: log-probability of the answer letter
(standard error about 2.4 points). Verdict: dead end. At equal or better text scores, the run s1-text62-lm1.0
(same speed, 62 % text, no real-speech TTS) reaches 0.12; the second step was therefore started from that run
instead.
Run stage_b_var: "voice" step, backbone slowed down
| Start | weights of s1-text62-lm1.0: base model + 2,600 steps with every parameter at 1e-4 on anchor_v5 plus a third pass of the text (62 % text-only examples) |
| Corpus | stage_b = anchor_v5 + a French numbers set seen twice per epoch (210,307 training examples) |
| Numbers set | 3,000 template sentences with dates, times, prices, percentages, addresses, years and quantities written in digits, read by the Qwen3-TTS fine-tuned on SIWIS and kept when Whisper's transcription matched (1,761 sentences); added because numbers were often read in English |
| Learning rate | 1e-4 for the audio modules and depth transformer, 1e-5 for the language-model backbone (lm_lr_scale 0.1), 300 warm-up steps, cosine decay |
| Batch | 8 x 8 accumulation = 64 examples, bf16, AdamW, gradient clipping 1.0 |
| Steps | 2 epochs = 6,574 steps; checkpoint_0 = step 3,287 (end of epoch 1), checkpoint_1 = final = step 6,574 |
Why the backbone is slowed down: in LFM2-Audio the Mimi codebook 0 (which sound comes next) is predicted from the backbone hidden state alone, so French pronunciation is learnt by the backbone, and the text knowledge lives in the same layers. Earlier runs on the same data showed that a backbone at full speed learns French pronunciation but loses text knowledge, while a backbone at scale 0.1 keeps the text. This step tests whether a slow second step keeps the text where the first step left it while polishing the voice.
Final evaluation (50 French and 50 English probe sentences, D42 sampling, mean of seeds 0, 1 and 2):
| Base | s1-text62-lm1.0 (start) |
stage_b_var final |
|
|---|---|---|---|
| French TTS WER (SIWIS prompt) | 0.695 | 0.122 (seed 0) | 0.100 ± 0.011 |
| French interleaved TTS WER (audio) | 2.56 | 0.124, text end 98.7 %, no loop | |
| English TTS WER (US female prompt) | 0.034 | 0.036 | |
| English interleaved TTS WER (audio) | 0.041, text end 100 % | ||
| DistillMOS, French TTS | 4.12 | 4.06 | |
| INCLUDE-fr (419 questions) | 42.2 % | 36.0 % | 35.8 % |
| Belebele-fr (900 questions) | 49.2 % | 39.6 % | 38.9 % |
Greedy decoding (indicative): French TTS 0.141, English TTS 0.032. At equal seed (0) the per-sentence mean French TTS WER is 0.123 against 0.132 for the starting point, and the 20 sentences probed at every validation moved from 0.156 to 0.159: within noise. Verdict: the slow step neither degraded nor recovered the text, and did not change the measured pronunciation; whether it sounds better (numbers, liaisons) is being judged by listening.
Training code and licence
Training code: finetuning/ sub-project of the project repository (commit 8c8ef9b) on top of the
liquid-audio fork with patches 0001-0013 (5f12fd7); lm_lr_scale is patch 0013. These weights are derivatives
of LFM2.5-Audio-1.5B and are distributed under its licence (see LICENSE). Training data licences are those of the
datasets listed above.
Résumé en français
Points intermédiaires de deux essais d'ancrage du français (plan v3, 8-9 octobre 2026), publiés pour libérer le
disque de la machine d'entraînement. stage_a : tout le modèle à 1e-4, 50 000 clips réels Emilia en cible TTS sous
un prompt générique, 2 000 pas ; impasse (TTS FR 0,19-0,22, INCLUDE 37,7 %). stage_b_var : départ des poids de
s1-text62-lm1.0, backbone à ×0,1, 2 époques avec un jeu de nombres en voix SIWIS ; TTS FR 0,100, texte inchangé
(INCLUDE 35,8 %, Belebele 38,9 %), écart avec le point de départ dans le bruit, verdict à l'écoute. Poids seuls
(pas d'état d'optimiseur) ; décodage recommandé : audio T 1,0 / top-k 4, texte en greedy.
Run s2_siwis_r1 (phase 2, plan v3, octobre 2026)
Notre recette de phase 2 (plan v3) : départ de l'étape B (stage_b_var/final) ; conversations v1 + outils v2 et v3, assistante en voix SIWIS (FR) et US female (EN) ; ce que dit l'assistante écrit en forme parlée (nombres en mots) ; dialogues à outils aussi en version texte seul (l'utilisateur écrit, l'assistante parle) ; rappel 25 % tiré de notre ancrage (ASR réel, TTS dans les voix de l'assistante) et ~10 % de texte rejoué ; tout le modèle à 1e-5, 3 époques.
Départ data/finetune/runs/stage_b_var/final, corpus s2_siwis_r1, lr 1e-05,
lm_lr_scale 1.0, 3.0 époque(s), 5670 pas, lot 1 × 16,
contexte 2048. Configuration complète : runs/s2_siwis_r1/params.json. Évaluation de fin (protocole
de phase 2 : 50 conversations et 50 scénarios d'outils v2 par langue, sondes jamais entraînées) : runs/s2_siwis_r1/eval/,
analyse des erreurs : runs/s2_siwis_r1/analyse_erreurs.txt, page d'écoute : listening/ (télécharger le dossier,
ouvrir index.html).
| Mesure (moyenne des graines quand il y en a plusieurs) | Valeur |
|---|---|
| TTS FR (WER, tirage D42) | 0.123 |
| TTS EN | 0.035 |
| Entrelacée FR / EN (WER audio) | 0.133 / 0.037 |
| Conversation FR : cohérence voix/texte | 0.102 |
| Conversation FR : langue correcte / fin de texte | 1.000 / 0.993 |
| Conversation FR : reprise de l'utilisateur | 0.000 |
| Conversation EN : cohérence voix/texte | 0.024 |
| Outils FR : décision correcte (appels) / bon outil | 0.945 / 0.918 |
| Outils FR : arguments identité / non-appels corrects | 0.890 / 0.964 |
| Outils EN : décision correcte (appels) / bon outil | 0.945 / 0.890 |
| Outils EN : arguments identité / non-appels corrects | 0.767 / 0.964 |
| Banc texte belebele_s2_siwis_r1_loglik | 0.424 |
| Banc texte include_s2_siwis_r1_loglik | 0.380 |
Run s2_siwis_r2 (phase 2, plan v3, octobre 2026)
R1 suivi d'une passe courte (0,5 époque, lr 5e-6) où les scénarios d'entraînement des cellules (famille / sorte) ratées par R1 sur les sondes sont sur-représentés x3, en audio et en texte seul ; les sondes elles-mêmes ne sont jamais entraînées.
Départ data/finetune/runs/s2_siwis_r1/final, corpus s2_siwis_r2, lr 5e-06,
lm_lr_scale 1.0, 0.5 époque(s), 945 pas, lot 1 × 16,
contexte 2048. Configuration complète : runs/s2_siwis_r2/params.json. Évaluation de fin (protocole
de phase 2 : 50 conversations et 50 scénarios d'outils v2 par langue, sondes jamais entraînées) : runs/s2_siwis_r2/eval/,
analyse des erreurs : runs/s2_siwis_r2/analyse_erreurs.txt, page d'écoute : listening/ (télécharger le dossier,
ouvrir index.html).
| Mesure (moyenne des graines quand il y en a plusieurs) | Valeur |
|---|---|
| TTS FR (WER, tirage D42) | 0.116 |
| TTS EN | 0.034 |
| Entrelacée FR / EN (WER audio) | 0.111 / 0.033 |
| Conversation FR : cohérence voix/texte | 0.132 |
| Conversation FR : langue correcte / fin de texte | 1.000 / 0.993 |
| Conversation FR : reprise de l'utilisateur | 0.000 |
| Conversation EN : cohérence voix/texte | 0.019 |
| Outils FR : décision correcte (appels) / bon outil | 0.904 / 0.877 |
| Outils FR : arguments identité / non-appels corrects | 0.781 / 0.988 |
| Outils EN : décision correcte (appels) / bon outil | 0.904 / 0.849 |
| Outils EN : arguments identité / non-appels corrects | 0.753 / 0.941 |
| Banc texte belebele_s2_siwis_r2_loglik | 0.423 |
| Banc texte include_s2_siwis_r2_loglik | 0.372 |