Labari Voice
AI & ML interests
None defined yet.
Recent Activity
Labari Voice
Speech datasets · low-resource languages
The speech data your models have never heard.
20+ languages and local varieties. 500M+ speakers covered. Native-speaker recordings, transcription, 14 annotation layers, commercial rights included: produced in studio, with native speakers under contract.
labari.dev · sales@labari.dev
Why this data doesn't exist
Speech models were trained on the languages that were easiest to collect. More than half of humanity lives in regions whose languages fall outside that set.
This is not a demand problem. It is a production problem, and it is why the gap has stayed open.
You cannot scrape this data. It has to be recorded: in country, in studio, with native speakers under contract, on terms that clear commercial rights at the source. Then transcribed, in languages whose written standards are still consolidating and whose competent transcribers are few.
That barrier is the reason the data isn't available. It is also what we built.
Languages
| Language | Locale |
|---|---|
| Fon | fon-BJ |
| Wolof | wo-SN |
| Beninese French | fr-BJ |
| Senegalese French | fr-SN |
| Ewe | ee-BJ |
| Yoruba | yo-BJ |
| Hausa | ha-BJ |
| Pulaar | fuc-SN |
| Bambara | bm-SN |
| Maninka | mlq-SN |
| Darija | ary-MA |
| Moroccan French | fr-MA |
| Swahili | sw-TZ |
| Tanzanian English | en-TZ |
| Dioula | dyu-CI |
| Baoulé | bci-CI |
| Ivorian French | fr-CI |
| Fulfulde | fub-CM |
| Pidgin | wes-CM |
| Cameroonian French | fr-CM |
Need a language that isn't listed? Our sourcing network extends beyond the published catalog.
What a Labari dataset contains
- Native-speaker recordings: studio-recorded, speakers under contract, documented recording chain.
- Transcription, with orthographic validation for languages whose written standards are still consolidating.
- 14 annotation layers: phonetics, semantics, timing and context.
- Quality reported, not claimed: gold-standard multi-annotator protocol, with Krippendorff α, WER and CER in the datasheet.
- Commercial rights secured at the source, GDPR compliant.
How we produce
End-to-end, in-house: speaker sourcing, studio recording, segmentation, transcription, orthographic validation and QA.
Every corpus carries its own audit trail: segmentation thresholds, per-segment quality measurements, SHA-256 checksums for sources and segments, and the full production parameters. Where a recording chain applies processing that affects a measurement, we say so and flag what it affects. We would rather publish a caveat than let a buyer find it after delivery.
That traceability is the product as much as the audio: it is what lets you filter a corpus on quality before training, and what lets you defend its provenance afterwards.
Licensing
- Catalog license: annual subscription access to one or more languages.
- Custom datasets: specify language, domain, register, speech type, speaker profile, volume and recording conditions; we produce to that brief.
- Temporary exclusivity: an exclusivity window on a language, domain or corpus, before it joins the shared catalog.
Public samples
The corpora published on this page are pilot releases: small, unlabeled, and intended for method review rather than training. They exist so the segmentation, measurement and documentation behind the catalog can be audited before you evaluate the annotated product.
The commercial catalog (transcribed, annotated, licensed) is not distributed here.