LFM2.5-Audio, French anchoring: experiment checkpoints (plan v3, October 2026)

Intermediate checkpoints of two fine-tuning experiments of LiquidAI/LFM2.5-Audio-1.5B towards a French voice assistant (French in the SIWIS voice, English in the base model's "US female" voice, while preserving the text abilities of the language model). They are experiment artefacts, not a release: they are published so that the training machine's disk can be freed without losing the weights, and so that any of them can serve as the starting point of a new run.

Layout

runs/<run>/checkpoint_N/model.safetensors   weights at an intermediate step (+ config.json of the base model)
runs/<run>/final/model.safetensors          weights at the last step
runs/<run>/params.json                      full training configuration written by the training tool
runs/<run>/eval/...                         metrics (summary.json, per-sentence CSV), no audio

Only the model weights are kept. Optimizer, scheduler and RNG states were not uploaded, so a run cannot be resumed exactly; a new run can start from any of these weights. The tokenizer, processor and audio detokenizer are those of the base model (LiquidAI/LFM2.5-Audio-1.5B@c362a0625dfe45aa588dce5f0ada28a7e5707628): download one folder, then either pass it to s2s-finetune --model-id <folder> in the training repository, or copy the base model's other files next to model.safetensors and load the folder with LFM2AudioModel.from_pretrained.

Recommended decoding (decision D42 of the project): sample the audio with temperature 1.0 and top-k 4, in sequential and interleaved generation alike; keep the text greedy.

Common data

Both runs build on the corpus anchor_v5 (207,033 training examples, context 1,024 tokens). Every example supervises the assistant turn only.

  • ASR on real multi-speaker speech: French from Emilia-YODAS, Multilingual LibriSpeech, Common Voice, VoxPopuli and CEFC-Orféo (31,056 clips); English from LibriTTS-R (21,211 clips). Prompts Perform ASR. The audio is in French. / Perform ASR. The audio is in English.
  • TTS and interleaved TTS (6 text tokens / 12 audio frames) on the assistant voice only: SIWIS studio recordings (9,750 clips, 10.7 h; Rcarvalo/siwis-fr-24k) and the assistant replies of synthetic dialogues re-synthesised in the target voice (French: 21,267 clips / 50.6 h with a Qwen3-TTS fine-tuned on SIWIS; English: 21,216 clips / 48.4 h with the base model's US female voice). Prompts Perform TTS. Use the French female voice., Perform TTS with interleaved text and audio. Use the French female voice. (and ... the US female voice. in English).
  • Text only (about half of the examples): 39,198 French and 14,727 English question/answer exchanges (Rcarvalo/text-replay-fr-en), seen twice, prompts Tu es un assistant utile. Réponds en français, de façon claire et exacte. / You are a helpful assistant. Answer in English, clearly and accurately.

Probe sentences and the validation/test dialogues of the next phase are excluded; splits are grouped by recording or dialogue.

Run stage_a: "language" step, whole model at full speed

Start base model LiquidAI/LFM2.5-Audio-1.5B@c362a06
Corpus anchor_v5 + 50,000 real French Emilia-YODAS clips used as TTS / interleaved TTS targets under the generic prompt Perform TTS. Use a French voice. (many speakers; SIWIS keeps its named prompt)
Learning rate 1e-4 for every parameter (lm_lr_scale 1.0), 150 warm-up steps, cosine decay
Batch 4 x 16 accumulation = 64 examples, bf16, AdamW, gradient clipping 1.0
Steps 2,000; checkpoints every 500 steps (checkpoint_0 = step 500 ... checkpoint_3 = step 2,000 = final)

Goal: learn French phonetics quickly from real multi-speaker speech before a second, slow step for the voice.

Checkpoint (step) INCLUDE-fr (419 q.) French TTS WER
base model 42.2 % 0.62
checkpoint_0 (500) 39.1 % 0.319
checkpoint_1 (1,000) 33.7 % 0.262
checkpoint_2 (1,500) 37.7 % 0.190
checkpoint_3 (2,000) 37.7 % 0.217

TTS: 50 fixed French probe sentences never seen in training, D42 sampling, seed 0, generic prompt Perform TTS. Use a French voice., transcribed by Whisper-small. INCLUDE: log-probability of the answer letter (standard error about 2.4 points). Verdict: dead end. At equal or better text scores, the run s1-text62-lm1.0 (same speed, 62 % text, no real-speech TTS) reaches 0.12; the second step was therefore started from that run instead.

Run stage_b_var: "voice" step, backbone slowed down

Start weights of s1-text62-lm1.0: base model + 2,600 steps with every parameter at 1e-4 on anchor_v5 plus a third pass of the text (62 % text-only examples)
Corpus stage_b = anchor_v5 + a French numbers set seen twice per epoch (210,307 training examples)
Numbers set 3,000 template sentences with dates, times, prices, percentages, addresses, years and quantities written in digits, read by the Qwen3-TTS fine-tuned on SIWIS and kept when Whisper's transcription matched (1,761 sentences); added because numbers were often read in English
Learning rate 1e-4 for the audio modules and depth transformer, 1e-5 for the language-model backbone (lm_lr_scale 0.1), 300 warm-up steps, cosine decay
Batch 8 x 8 accumulation = 64 examples, bf16, AdamW, gradient clipping 1.0
Steps 2 epochs = 6,574 steps; checkpoint_0 = step 3,287 (end of epoch 1), checkpoint_1 = final = step 6,574

Why the backbone is slowed down: in LFM2-Audio the Mimi codebook 0 (which sound comes next) is predicted from the backbone hidden state alone, so French pronunciation is learnt by the backbone, and the text knowledge lives in the same layers. Earlier runs on the same data showed that a backbone at full speed learns French pronunciation but loses text knowledge, while a backbone at scale 0.1 keeps the text. This step tests whether a slow second step keeps the text where the first step left it while polishing the voice.

Final evaluation (50 French and 50 English probe sentences, D42 sampling, mean of seeds 0, 1 and 2):

Base s1-text62-lm1.0 (start) stage_b_var final
French TTS WER (SIWIS prompt) 0.695 0.122 (seed 0) 0.100 ± 0.011
French interleaved TTS WER (audio) 2.56 0.124, text end 98.7 %, no loop
English TTS WER (US female prompt) 0.034 0.036
English interleaved TTS WER (audio) 0.041, text end 100 %
DistillMOS, French TTS 4.12 4.06
INCLUDE-fr (419 questions) 42.2 % 36.0 % 35.8 %
Belebele-fr (900 questions) 49.2 % 39.6 % 38.9 %

Greedy decoding (indicative): French TTS 0.141, English TTS 0.032. At equal seed (0) the per-sentence mean French TTS WER is 0.123 against 0.132 for the starting point, and the 20 sentences probed at every validation moved from 0.156 to 0.159: within noise. Verdict: the slow step neither degraded nor recovered the text, and did not change the measured pronunciation; whether it sounds better (numbers, liaisons) is being judged by listening.

Training code and licence

Training code: finetuning/ sub-project of the project repository (commit 8c8ef9b) on top of the liquid-audio fork with patches 0001-0013 (5f12fd7); lm_lr_scale is patch 0013. These weights are derivatives of LFM2.5-Audio-1.5B and are distributed under its licence (see LICENSE). Training data licences are those of the datasets listed above.

Résumé en français

Points intermédiaires de deux essais d'ancrage du français (plan v3, 8-9 octobre 2026), publiés pour libérer le disque de la machine d'entraînement. stage_a : tout le modèle à 1e-4, 50 000 clips réels Emilia en cible TTS sous un prompt générique, 2 000 pas ; impasse (TTS FR 0,19-0,22, INCLUDE 37,7 %). stage_b_var : départ des poids de s1-text62-lm1.0, backbone à ×0,1, 2 époques avec un jeu de nombres en voix SIWIS ; TTS FR 0,100, texte inchangé (INCLUDE 35,8 %, Belebele 38,9 %), écart avec le point de départ dans le bruit, verdict à l'écoute. Poids seuls (pas d'état d'optimiseur) ; décodage recommandé : audio T 1,0 / top-k 4, texte en greedy.

Run s2_siwis_r1 (phase 2, plan v3, octobre 2026)

Notre recette de phase 2 (plan v3) : départ de l'étape B (stage_b_var/final) ; conversations v1 + outils v2 et v3, assistante en voix SIWIS (FR) et US female (EN) ; ce que dit l'assistante écrit en forme parlée (nombres en mots) ; dialogues à outils aussi en version texte seul (l'utilisateur écrit, l'assistante parle) ; rappel 25 % tiré de notre ancrage (ASR réel, TTS dans les voix de l'assistante) et ~10 % de texte rejoué ; tout le modèle à 1e-5, 3 époques.

Départ data/finetune/runs/stage_b_var/final, corpus s2_siwis_r1, lr 1e-05, lm_lr_scale 1.0, 3.0 époque(s), 5670 pas, lot 1 × 16, contexte 2048. Configuration complète : runs/s2_siwis_r1/params.json. Évaluation de fin (protocole de phase 2 : 50 conversations et 50 scénarios d'outils v2 par langue, sondes jamais entraînées) : runs/s2_siwis_r1/eval/, analyse des erreurs : runs/s2_siwis_r1/analyse_erreurs.txt, page d'écoute : listening/ (télécharger le dossier, ouvrir index.html).

Mesure (moyenne des graines quand il y en a plusieurs) Valeur
TTS FR (WER, tirage D42) 0.123
TTS EN 0.035
Entrelacée FR / EN (WER audio) 0.133 / 0.037
Conversation FR : cohérence voix/texte 0.102
Conversation FR : langue correcte / fin de texte 1.000 / 0.993
Conversation FR : reprise de l'utilisateur 0.000
Conversation EN : cohérence voix/texte 0.024
Outils FR : décision correcte (appels) / bon outil 0.945 / 0.918
Outils FR : arguments identité / non-appels corrects 0.890 / 0.964
Outils EN : décision correcte (appels) / bon outil 0.945 / 0.890
Outils EN : arguments identité / non-appels corrects 0.767 / 0.964
Banc texte belebele_s2_siwis_r1_loglik 0.424
Banc texte include_s2_siwis_r1_loglik 0.380

Run s2_siwis_r2 (phase 2, plan v3, octobre 2026)

R1 suivi d'une passe courte (0,5 époque, lr 5e-6) où les scénarios d'entraînement des cellules (famille / sorte) ratées par R1 sur les sondes sont sur-représentés x3, en audio et en texte seul ; les sondes elles-mêmes ne sont jamais entraînées.

Départ data/finetune/runs/s2_siwis_r1/final, corpus s2_siwis_r2, lr 5e-06, lm_lr_scale 1.0, 0.5 époque(s), 945 pas, lot 1 × 16, contexte 2048. Configuration complète : runs/s2_siwis_r2/params.json. Évaluation de fin (protocole de phase 2 : 50 conversations et 50 scénarios d'outils v2 par langue, sondes jamais entraînées) : runs/s2_siwis_r2/eval/, analyse des erreurs : runs/s2_siwis_r2/analyse_erreurs.txt, page d'écoute : listening/ (télécharger le dossier, ouvrir index.html).

Mesure (moyenne des graines quand il y en a plusieurs) Valeur
TTS FR (WER, tirage D42) 0.116
TTS EN 0.034
Entrelacée FR / EN (WER audio) 0.111 / 0.033
Conversation FR : cohérence voix/texte 0.132
Conversation FR : langue correcte / fin de texte 1.000 / 0.993
Conversation FR : reprise de l'utilisateur 0.000
Conversation EN : cohérence voix/texte 0.019
Outils FR : décision correcte (appels) / bon outil 0.904 / 0.877
Outils FR : arguments identité / non-appels corrects 0.781 / 0.988
Outils EN : décision correcte (appels) / bon outil 0.904 / 0.849
Outils EN : arguments identité / non-appels corrects 0.753 / 0.941
Banc texte belebele_s2_siwis_r2_loglik 0.423
Banc texte include_s2_siwis_r2_loglik 0.372
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rcarvalo/lfm2a-fr-v3-checkpoints

Finetuned
(8)
this model