Whistle ONNX

ONNX export of Cactus-Compute/whistle, built from a PyTorch re-implementation that loads the deployed, dequantized weights from whistle.cact. On the reference clips (en/de/fr/es) the output matches the Cactus needle engine exactly:

check result
transcript text (greedy) identical on 4/4 clips
language ID (auto) correct on 4/4
log-mel features vs engine max abs error ≤ 1.1e-4
word timestamps (DTW, 80 ms frames) 85/92 boundaries identical, the rest ±1 frame
word probabilities max abs diff 0.055

Live demo (WebGPU, in the browser): mrfakename/whistle-webgpu

Files

file
onnx/encoder.onnx (+.data) log-mel [T,80] → cross_k [8,8,S,48], cross_v [8,8,S,64] (fp32, 72 MB)
onnx/decoder.onnx (+.data) one decoder step with KV cache (fp32, 73 MB)
onnx/*_fp16.onnx fp16 weights/compute (numerically sensitive ops kept in fp32)
whistle.pack 17 MB pack: original 2/3/4-bit CQ weights + fp16/fp32 rest, tokenizer, mel filterbank, engram tables. js/pack.js rebuilds any of the .onnx.data files from it byte-exactly (fp32)
js/ reference pipeline: mel features, BPE tokenizer, engram lookup, greedy decoding with LID, keyword biasing, DTW timestamps (whistle.js), onnxruntime backend (backend.js)
scripts/ PyTorch model (wt.py), weight extraction (weights.py), export (export_onnx.py), fp16 conversion (to_fp16.py), pack builder (pack.py)

Decoder step inputs: token [1], pos [1], efeat [2,4,512] (engram features, computed on the host from the token history, see Engram in whistle.js), prev [8,2,608], past_k [8,2,P,48], past_v [8,2,P,64], cross_k, cross_v. Outputs: logits [8199], prev_out, present_k/v (P+1), xattn [4,8,S] (cross-attention of layers 4-7, for timestamps).

Usage (Node)

npm i onnxruntime-node
node js/example.mjs audio.wav            # 16 kHz mono 16-bit WAV
node js/example.mjs audio.wav altavoz    # with keyword biasing

In the browser the same js/ files run on onnxruntime-web. The demo puts the encoder on the WebGPU EP (fp16) and the decoder on WASM, which is faster than WebGPU for single-token steps.

Notes

  • Decoding is greedy; keyword biasing (+2 logit on the first token, +5 on continuations) approximates the engine's.
  • The encoder's Sinkhorn uses a max-shifted log-sum-exp on its first iteration, because onnxruntime-web's WebGPU ReduceLogSumExp is not overflow-safe. This doesn't change the CPU results.
  • To rebuild: put whistle.cact and whistle.safetensors from the base repo next to the scripts, then run python weights.py && python export_onnx.py && python to_fp16.py onnx onnx16 && python pack.py release.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mrfakename/whistle-ONNX

Quantized
(3)
this model

Space using mrfakename/whistle-ONNX 1