O1-Sound
A greeting wake-word spotter built on the O-Series liquid core.
Always-on keyword spotting on a battery: the model listens continuously, holds a fixed-size state no matter how long the microphone has been open, and fires when it hears a greeting — "hello", "hola", "bonjour", "你好".
mic → log-mel (fixed, no weights) → liquid core (multi-timescale τ) → wake / not-wake
↑ carried state: 5,120 bytes, constant
Files
| File | Description |
|---|---|
o1sound.pt |
Binary wake/not-wake research checkpoint (Run 7, 2026-08-18), 4.8 MB |
o1sound_multiclass.pt |
Multiclass experiment checkpoint (2026-09), one class per greeting + "other" |
o1sound.onnx |
Streaming ONNX graph (fp32, single-frame step()), 5.04 MB |
o1sound.metrics.json / o1sound_multiclass.metrics.json |
Per-language metrics for each checkpoint |
o1sound_multiclass_v2.pt |
Multiclass v2 checkpoint (Run 8, 2026-09-29), 9.8 MB — hidden 896, 40 epochs, 400 negatives |
o1sound_neg_v1.pt |
Binary v3 checkpoint, real negative audio (Run 9, 2026-10-04), 9.8 MB — FAR ≤ 1% on the held-out test (0.97% @ thr 0.9906) |
o1sound_neg_v2.pt |
v3 full-negative retrain (9,706 clips) — honest negative result (class imbalance broke recall; dev balanced 0.732) |
o1sound_neg_v1.metrics.json / o1sound_neg_v2.metrics.json |
Dev metrics for the v3 checkpoints |
test.json / test_multiclass.json / test_v2.json / test_neg_v1_exact.json |
Held-out test results (Run 7 / multiclass / Run 8 / Run 9 exact sweep) |
Honest status — what is and isn't validated
Validated (runs today, in CI)
- 32 tests pass: streaming
step()is numerically identical to the batchedforward(), carried state stays constant over 500 frames. - ONNX export: 5.03 MB fp32, 1.27 MB int8, max |onnx − torch| = 3.7e-09.
- Carried state: 5,120 bytes per stream, independent of stream length — a property of the recurrence, verified by test.
Measured — the architecture works, the multilingual framing does not (yet)
- Run 2 (2026-08-01, English only, 301 wake clips): FRR 0.146 at FAR 0.046 on held-out test; dev accuracy 0.919 against a 0.780 never-fire baseline.
- Run 1 (2026-07-31, 9 languages, 91 wake clips): FRR 0.909 at FAR 0.049 — worse than saying no unconditionally at that data scale.
- Run 7 (2026-08-18, 10 languages, full MSWC): dev acc 0.885 (balanced 0.849); English is the only language with meaningful data (FRR 0.098 at FAR 0.133, 41 clips). Worst-language FRR is 1.000 (2 positive clips) — the multilingual claim is bounded by per-language data, not the architecture.
- Multiclass experiment (2026-09, 10 languages, OR-head): dev balanced accuracy 0.84; test FRR 0.35 @ FAR 0.046 overall — multilingual remains unsupported at this data scale.
- Run 8 (v2, 2026-09-29, GPU A100) — improved recipe (hidden 896, 40
epochs, 400 negatives, multiclass OR inference): dev acc 0.838, dev
balanced 0.8815, worst-language dev FRR 0.600 (sv-SE; Run 7:
1.000). On the held-out test split no operating point reaches FAR ≤ 1%
— the production bar was still unmet at this point. Checkpoint:
o1sound_multiclass_v2.pt(9.8 MB, in this repo). - Run 9 (v3, 2026-10-04, real negative audio) — Speech Commands extras
- full MSWC negative pool (3,000 balanced clips in training): dev balanced
0.8913 (Run 8: 0.8815), worst-language dev FRR 0.500 (Run 8:
0.600). On the held-out test set (167 pos / 1,859 neg) the FAR ≤ 1% bar
is crossed: FAR 0.97% (18/1,859) at thr 0.9906 — once the threshold
sweep is exact (candidates = all unique scores; the earlier "1.02%
minimum" was grid quantization, not a model limit). At thr 0.9994, FAR
reaches 0 (0/1,859) with FRR 0.81. Honest note: the full 9,706-clip
retrain (60 epochs) is a negative result — class imbalance (~1:10)
collapsed dev balanced accuracy to 0.7319; the balanced 3,000-clip sample
(v3) is the best configuration. Checkpoints:
o1sound_neg_v1.pt(best),o1sound_neg_v2.pt(kept for honesty). The multilingual claim remains bounded by per-language data (worst-language test FRR 1.000 at 2–3 clips), independent of this FAR fix.
- full MSWC negative pool (3,000 balanced clips in training): dev balanced
0.8913 (Run 8: 0.8815), worst-language dev FRR 0.500 (Run 8:
0.600). On the held-out test set (167 pos / 1,859 neg) the FAR ≤ 1% bar
is crossed: FAR 0.97% (18/1,859) at thr 0.9906 — once the threshold
sweep is exact (candidates = all unique scores; the earlier "1.02%
minimum" was grid quantization, not a model limit). At thr 0.9994, FAR
reaches 0 (0/1,859) with FRR 0.81. Honest note: the full 9,706-clip
retrain (60 epochs) is a negative result — class imbalance (~1:10)
collapsed dev balanced accuracy to 0.7319; the balanced 3,000-clip sample
(v3) is the best configuration. Checkpoints:
Not production-grade. Deployed wake words run single-digit FRR at a
false-accept rate quoted per hour, not per clip. Latency and power are
unmeasured on real hardware (the Python step() loop runs ~0.5 ms/frame on a
desktop CPU). On out-of-domain audio (e.g. pure tones) it can false-trigger.
Run 9's FAR ≤ 1% is a per-clip held-out test figure; the per-hour field
equivalent is not measured.
Model sizes (measured, n_classes=2)
| hidden | layers | params | fp32 | int8 | carried state |
|---|---|---|---|---|---|
| 384 | 2 | 483,984 | 1.94 MB | 0.48 MB | 3,072 B |
| 512 | 2 | 841,872 | 3.37 MB | 0.84 MB | 4,096 B |
| 640 | 2 | 1,298,064 | 5.19 MB | 1.30 MB | 5,120 B |
| 512 | 3 | 1,367,184 | 5.47 MB | 1.37 MB | 6,144 B |
| 768 | 2 | 1,852,560 | 7.41 MB | 1.85 MB | 6,144 B |
hidden=640, layers=2 is the default: the largest 2-layer configuration that
still clears 7 MB in fp32 once ONNX graph overhead is counted.
Usage
The exported ONNX graph is the single-frame step() — one frame in, logits and
the next state out. Full training / export / streaming-demo code lives in the
GitHub repository:
- Code, training and streaming demo: AwareLiquid/O1-Sound
- Research release notes: release
research-2026-08-18
Why a liquid core
Each channel carries its own learnable time constant τ, parameterised as
softplus(log_tau) + tau_min and initialised geometrically across 10–240 ms.
Short-τ channels track the current phoneme; long-τ channels hold the envelope of
the whole word — separating a wake phrase from a phonetically close neighbour
without stacking depth. Streaming is O(1): the state does not grow with how long
the microphone has been open.
Related
- AwareLiquid/M1 — the MT-LNN / O-Series research line this core comes from
- AwareLiquid/M2 — next-generation architecture research
- awareliquid.ai — benchmarks and retractions
Try it
- 🎮 Interactive demo (record or upload audio, runs in your browser): https://huggingface.co/spaces/AwareLiquid/O1-Sound-demo
License
MIT.