O1-Sound

A greeting wake-word spotter built on the O-Series liquid core.

Always-on keyword spotting on a battery: the model listens continuously, holds a fixed-size state no matter how long the microphone has been open, and fires when it hears a greeting — "hello", "hola", "bonjour", "你好".

mic → log-mel (fixed, no weights) → liquid core (multi-timescale τ) → wake / not-wake
                                     ↑ carried state: 5,120 bytes, constant

Files

File Description
o1sound.pt Binary wake/not-wake research checkpoint (Run 7, 2026-08-18), 4.8 MB
o1sound_multiclass.pt Multiclass experiment checkpoint (2026-09), one class per greeting + "other"
o1sound.onnx Streaming ONNX graph (fp32, single-frame step()), 5.04 MB
o1sound.metrics.json / o1sound_multiclass.metrics.json Per-language metrics for each checkpoint
o1sound_multiclass_v2.pt Multiclass v2 checkpoint (Run 8, 2026-09-29), 9.8 MB — hidden 896, 40 epochs, 400 negatives
o1sound_neg_v1.pt Binary v3 checkpoint, real negative audio (Run 9, 2026-10-04), 9.8 MB — FAR ≤ 1% on the held-out test (0.97% @ thr 0.9906)
o1sound_neg_v2.pt v3 full-negative retrain (9,706 clips) — honest negative result (class imbalance broke recall; dev balanced 0.732)
o1sound_neg_v1.metrics.json / o1sound_neg_v2.metrics.json Dev metrics for the v3 checkpoints
test.json / test_multiclass.json / test_v2.json / test_neg_v1_exact.json Held-out test results (Run 7 / multiclass / Run 8 / Run 9 exact sweep)

Honest status — what is and isn't validated

Validated (runs today, in CI)

  • 32 tests pass: streaming step() is numerically identical to the batched forward(), carried state stays constant over 500 frames.
  • ONNX export: 5.03 MB fp32, 1.27 MB int8, max |onnx − torch| = 3.7e-09.
  • Carried state: 5,120 bytes per stream, independent of stream length — a property of the recurrence, verified by test.

Measured — the architecture works, the multilingual framing does not (yet)

  • Run 2 (2026-08-01, English only, 301 wake clips): FRR 0.146 at FAR 0.046 on held-out test; dev accuracy 0.919 against a 0.780 never-fire baseline.
  • Run 1 (2026-07-31, 9 languages, 91 wake clips): FRR 0.909 at FAR 0.049 — worse than saying no unconditionally at that data scale.
  • Run 7 (2026-08-18, 10 languages, full MSWC): dev acc 0.885 (balanced 0.849); English is the only language with meaningful data (FRR 0.098 at FAR 0.133, 41 clips). Worst-language FRR is 1.000 (2 positive clips) — the multilingual claim is bounded by per-language data, not the architecture.
  • Multiclass experiment (2026-09, 10 languages, OR-head): dev balanced accuracy 0.84; test FRR 0.35 @ FAR 0.046 overall — multilingual remains unsupported at this data scale.
  • Run 8 (v2, 2026-09-29, GPU A100) — improved recipe (hidden 896, 40 epochs, 400 negatives, multiclass OR inference): dev acc 0.838, dev balanced 0.8815, worst-language dev FRR 0.600 (sv-SE; Run 7: 1.000). On the held-out test split no operating point reaches FAR ≤ 1% — the production bar was still unmet at this point. Checkpoint: o1sound_multiclass_v2.pt (9.8 MB, in this repo).
  • Run 9 (v3, 2026-10-04, real negative audio) — Speech Commands extras
    • full MSWC negative pool (3,000 balanced clips in training): dev balanced 0.8913 (Run 8: 0.8815), worst-language dev FRR 0.500 (Run 8: 0.600). On the held-out test set (167 pos / 1,859 neg) the FAR ≤ 1% bar is crossed: FAR 0.97% (18/1,859) at thr 0.9906 — once the threshold sweep is exact (candidates = all unique scores; the earlier "1.02% minimum" was grid quantization, not a model limit). At thr 0.9994, FAR reaches 0 (0/1,859) with FRR 0.81. Honest note: the full 9,706-clip retrain (60 epochs) is a negative result — class imbalance (~1:10) collapsed dev balanced accuracy to 0.7319; the balanced 3,000-clip sample (v3) is the best configuration. Checkpoints: o1sound_neg_v1.pt (best), o1sound_neg_v2.pt (kept for honesty). The multilingual claim remains bounded by per-language data (worst-language test FRR 1.000 at 2–3 clips), independent of this FAR fix.

Not production-grade. Deployed wake words run single-digit FRR at a false-accept rate quoted per hour, not per clip. Latency and power are unmeasured on real hardware (the Python step() loop runs ~0.5 ms/frame on a desktop CPU). On out-of-domain audio (e.g. pure tones) it can false-trigger. Run 9's FAR ≤ 1% is a per-clip held-out test figure; the per-hour field equivalent is not measured.

Model sizes (measured, n_classes=2)

hidden layers params fp32 int8 carried state
384 2 483,984 1.94 MB 0.48 MB 3,072 B
512 2 841,872 3.37 MB 0.84 MB 4,096 B
640 2 1,298,064 5.19 MB 1.30 MB 5,120 B
512 3 1,367,184 5.47 MB 1.37 MB 6,144 B
768 2 1,852,560 7.41 MB 1.85 MB 6,144 B

hidden=640, layers=2 is the default: the largest 2-layer configuration that still clears 7 MB in fp32 once ONNX graph overhead is counted.

Usage

The exported ONNX graph is the single-frame step() — one frame in, logits and the next state out. Full training / export / streaming-demo code lives in the GitHub repository:

Why a liquid core

Each channel carries its own learnable time constant τ, parameterised as softplus(log_tau) + tau_min and initialised geometrically across 10–240 ms. Short-τ channels track the current phoneme; long-τ channels hold the envelope of the whole word — separating a wake phrase from a phonetically close neighbour without stacking depth. Streaming is O(1): the state does not grow with how long the microphone has been open.

Related

Try it

License

MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using AwareLiquid/O1-Sound 1

Collection including AwareLiquid/O1-Sound