Auto-generated by Claude Sonnet 5.5 (claude-sonnet-5-5) via Claude Code — NOT human-reviewed.

Allosaurus uni2005 (ONNX)

This repository holds the acoustic model of Allosaurus, release uni2005, exported to ONNX. Allosaurus is a universal phone recogniser: a five-layer bidirectional LSTM over stacked MFCC frames, trained with CTC on a multilingual allophone system. It outputs phones, not words, for speech in any language. The graph covers the acoustic model only; the feature extraction and the CTC decoding run outside it.

Files

File Content
allosaurus_uni2005.onnx the acoustic model
allosaurus_reference.json the phones Allosaurus itself emits for 12 test clips, used for the parity check
conversion/export_allosaurus.py the export script

Tensors

The graph has one input and one output, both with a dynamic time axis.

Name Shape Type Meaning
feats (1, T, 120) float32 stacked MFCC frames
logits (1, T, 230) float32 per-frame scores over 229 phone units; unit 0 is the CTC blank

Preprocessing

The caller computes the features from 8 kHz mono audio, using the parameters of the model's pm_config.json:

  1. Scale samples in [-1, 1] to 16-bit sample values.
  2. Compute 40 Kaldi MFCCs per frame with a Povey window, 40 mel banks over 40 to 3800 Hz, a DCT-II, lifter 22 and no energy coefficient.
  3. Normalise each coefficient to zero mean and unit variance over the utterance.
  4. Stack each frame with its two neighbours, wrapping at the edges, giving 120 values per frame, and keep every third stacked frame.

Decoding is greedy CTC: take the arg-max unit per frame, merge repeats and drop the blank. The unit-to-phone table is that of uni2005 in the Allosaurus release.

Export

conversion/export_allosaurus.py loads the extracted uni2005 release (https://github.com/xinjli/allosaurus/releases/download/v1.0/latest.tar.gz) in an environment with allosaurus, torch and onnx, wraps the BLSTM without sequence packing for a single utterance, and writes the graph. With --clips it also writes allosaurus_reference.json: the phones Allosaurus's own recogniser emits for each test clip after the clip is resampled to 8 kHz and stored as 16-bit PCM.

Parity check

The exported graph, fed the features described above and decoded greedily, was compared with the phones in allosaurus_reference.json over 12 clips with 172 reference phones. Eleven clips are identical. One clip has one extra phone, 1 phone edit in 172. The difference falls on a frame of digital silence, where a small numerical difference between runtimes decides between the blank and a phone.

Licence and attribution

Allosaurus code and the pretrained uni2005 model are released under GPL-3.0 by their authors; this export is a derivative and carries the same licence (GPL-3.0). The original work is at xinjli/allosaurus.

If you use this model, cite:

@inproceedings{li2020universal,
  title={Universal Phone Recognition with a Multilingual Allophone System},
  author={Li, Xinjian and others},
  booktitle={ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing},
  year={2020}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support