Auto-generated by Claude Sonnet 5.5 (claude-sonnet-5-5) via Claude Code — NOT human-reviewed.
Allosaurus uni2005 (ONNX)
This repository holds the acoustic model of Allosaurus, release uni2005, exported to ONNX. Allosaurus is a universal phone recogniser: a five-layer bidirectional LSTM over stacked MFCC frames, trained with CTC on a multilingual allophone system. It outputs phones, not words, for speech in any language. The graph covers the acoustic model only; the feature extraction and the CTC decoding run outside it.
Files
| File | Content |
|---|---|
allosaurus_uni2005.onnx |
the acoustic model |
allosaurus_reference.json |
the phones Allosaurus itself emits for 12 test clips, used for the parity check |
conversion/export_allosaurus.py |
the export script |
Tensors
The graph has one input and one output, both with a dynamic time axis.
| Name | Shape | Type | Meaning |
|---|---|---|---|
feats |
(1, T, 120) |
float32 | stacked MFCC frames |
logits |
(1, T, 230) |
float32 | per-frame scores over 229 phone units; unit 0 is the CTC blank |
Preprocessing
The caller computes the features from 8 kHz mono audio, using the parameters of the model's pm_config.json:
- Scale samples in
[-1, 1]to 16-bit sample values. - Compute 40 Kaldi MFCCs per frame with a Povey window, 40 mel banks over 40 to 3800 Hz, a DCT-II, lifter 22 and no energy coefficient.
- Normalise each coefficient to zero mean and unit variance over the utterance.
- Stack each frame with its two neighbours, wrapping at the edges, giving 120 values per frame, and keep every third stacked frame.
Decoding is greedy CTC: take the arg-max unit per frame, merge repeats and drop the blank. The unit-to-phone table is that of uni2005 in the Allosaurus release.
Export
conversion/export_allosaurus.py loads the extracted uni2005 release (https://github.com/xinjli/allosaurus/releases/download/v1.0/latest.tar.gz) in an environment with allosaurus, torch and onnx, wraps the BLSTM without sequence packing for a single utterance, and writes the graph. With --clips it also writes allosaurus_reference.json: the phones Allosaurus's own recogniser emits for each test clip after the clip is resampled to 8 kHz and stored as 16-bit PCM.
Parity check
The exported graph, fed the features described above and decoded greedily, was compared with the phones in allosaurus_reference.json over 12 clips with 172 reference phones. Eleven clips are identical. One clip has one extra phone, 1 phone edit in 172. The difference falls on a frame of digital silence, where a small numerical difference between runtimes decides between the blank and a phone.
Licence and attribution
Allosaurus code and the pretrained uni2005 model are released under GPL-3.0 by their authors; this export is a derivative and carries the same licence (GPL-3.0). The original work is at xinjli/allosaurus.
If you use this model, cite:
@inproceedings{li2020universal,
title={Universal Phone Recognition with a Multilingual Allophone System},
author={Li, Xinjian and others},
booktitle={ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing},
year={2020}
}