ECAPA-TDNN VoxLingua107 (spoken language identification)
SpeechBrain's ECAPA-TDNN spoken-language identifier over VoxLingua107's 107 languages, exported for loom.cpp. Family 13: audio in, one language distribution per clip out.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from speechbrain/lang-id-voxlingua107-ecapa. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
apache-2.0, inherited from the base model above.
Language(s)
ab, af, am, ar, as, az, ba, be, bg, bi, bo, br, bs, ca, ceb, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fo, fr, gl, gn, gu, gv, ha, haw, hi, hr, ht, hu, hy, ia, id, is, it, he, ja, jv, ka, kk, km, kn, ko, la, lm, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sco, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ud, uz, vi, war, yi, yo, zh
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/ecapa-voxlingua107-loom")
# Audio is a mono float list at 16 kHz, one whole clip per call. The answer is one row: a probability
# for every language the model knows, labelled "<code>: <name>".
result = model.speech2class.infer(audio)
print(result.best[0])
# en: English
# Closely related languages share probability, so read the top few rather than only the first:
for label, p in result.top(3):
print(f"{label:24s} {p:.3f}")
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
One whole clip per call. The export takes no length, so a batch of padded clips cannot be passed: the padding would be heard as part of the audio. Call it once per clip.
Trained on VoxLingua107, which is speech scraped from YouTube: a few seconds of clean speech is plenty, and closely related languages (the upstream card's examples are Bosnian/Croatian/Serbian and the Scandinavian languages) are routinely confused -- read the top few answers rather than only the first. Audio must be mono 16 kHz.
Files
ecapa-voxlingua107.gguf-- the model, exported with loom-exporter.
- Downloads last month
- 19
We're not able to determine the quantization variants.
Model tree for loom-ai-org/ecapa-voxlingua107-loom
Base model
speechbrain/lang-id-voxlingua107-ecapa