Whistle ONNX
ONNX export of Cactus-Compute/whistle, built from a PyTorch
re-implementation that loads the deployed, dequantized weights from whistle.cact. On the reference clips (en/de/fr/es)
the output matches the Cactus needle engine exactly:
| check | result |
|---|---|
| transcript text (greedy) | identical on 4/4 clips |
| language ID (auto) | correct on 4/4 |
| log-mel features vs engine | max abs error ≤ 1.1e-4 |
| word timestamps (DTW, 80 ms frames) | 85/92 boundaries identical, the rest ±1 frame |
| word probabilities | max abs diff 0.055 |
Live demo (WebGPU, in the browser): mrfakename/whistle-webgpu
Files
| file | |
|---|---|
onnx/encoder.onnx (+.data) |
log-mel [T,80] → cross_k [8,8,S,48], cross_v [8,8,S,64] (fp32, 72 MB) |
onnx/decoder.onnx (+.data) |
one decoder step with KV cache (fp32, 73 MB) |
onnx/*_fp16.onnx |
fp16 weights/compute (numerically sensitive ops kept in fp32) |
whistle.pack |
17 MB pack: original 2/3/4-bit CQ weights + fp16/fp32 rest, tokenizer, mel filterbank, engram tables. js/pack.js rebuilds any of the .onnx.data files from it byte-exactly (fp32) |
js/ |
reference pipeline: mel features, BPE tokenizer, engram lookup, greedy decoding with LID, keyword biasing, DTW timestamps (whistle.js), onnxruntime backend (backend.js) |
scripts/ |
PyTorch model (wt.py), weight extraction (weights.py), export (export_onnx.py), fp16 conversion (to_fp16.py), pack builder (pack.py) |
Decoder step inputs: token [1], pos [1], efeat [2,4,512] (engram features, computed on the host from the
token history, see Engram in whistle.js), prev [8,2,608], past_k [8,2,P,48], past_v [8,2,P,64], cross_k, cross_v.
Outputs: logits [8199], prev_out, present_k/v (P+1), xattn [4,8,S] (cross-attention of layers 4-7, for timestamps).
Usage (Node)
npm i onnxruntime-node
node js/example.mjs audio.wav # 16 kHz mono 16-bit WAV
node js/example.mjs audio.wav altavoz # with keyword biasing
In the browser the same js/ files run on onnxruntime-web. The demo puts the encoder on the WebGPU EP (fp16) and the decoder
on WASM, which is faster than WebGPU for single-token steps.
Notes
- Decoding is greedy; keyword biasing (+2 logit on the first token, +5 on continuations) approximates the engine's.
- The encoder's Sinkhorn uses a max-shifted log-sum-exp on its first iteration, because onnxruntime-web's WebGPU
ReduceLogSumExpis not overflow-safe. This doesn't change the CPU results. - To rebuild: put
whistle.cactandwhistle.safetensorsfrom the base repo next to the scripts, then runpython weights.py && python export_onnx.py && python to_fp16.py onnx onnx16 && python pack.py release.
Model tree for mrfakename/whistle-ONNX
Base model
Cactus-Compute/whistle