pVAD v2: speaker-conditioned frame VAD for telephony barge-in
pVAD v2 labels every 10 ms of an 8 kHz caller stream as silence, the enrolled caller, or someone else. "Someone else" covers a bystander, a TV or laptop, or the agent's own echo. The model is conditioned on a 192-d speaker embedding of the caller.
v2 uses the same 171k-parameter graph and streaming contract as v1, so it is a drop-in replacement. The difference is the training data:
- v1 learned mostly near/far loudness.
- v2 was trained on 90k synthetic scenes built to separate who is speaking:
- the speakers: same- and other-gender bystanders, TV and laptop playback, agent echo, the caller on speakerphone, far rooms, music, babble, short words and handovers;
- the channels: G.711, AMR, Opus, G.722 and web audio;
- the voices: public multilingual speech plus caller voices from Telenow's own Indian call recordings.
No audio is released with this model.
Results
Synthetic (held-out speakers)
Each cell has 250 six-second mixtures of speakers unseen in training. Each model runs at its own caller-safe calibration (see below). The table gives the share of another voice's utterances that the model flags as not the caller.
| scenario | v1 | v2 |
|---|---|---|
| agent echo (caller silent / double-talk) | 1.0 % | 56.7 % |
| near-field person, other gender | 0 % | 48.6 % |
| bystander/TV only, and handovers | 0.9 % | 44.2 % |
| TV 1β3 m | 0 % | 40.5 % |
| laptop/phone 0.2β0.8 m | 0 % | 36.8 % |
| caller on speakerphone, bystander close | 0 % | 25.6 % |
| far room 1β4 m | 2.5 % | 23.9 % |
| near-field person, same gender | 0 % | 7.5 % |
| same gender, low pitch 130β180 Hz | 0 % | 8.8 % |
| enrollment 1.5β2 s, 20β40 % contaminated | 0 % | 4.0 % |
Caller utterances dropped: at most 1.2 % in any cell (caller-only cell: 0.4 %).
Real calls
On real calls v2 has not yet proven better than v1:
- Held-out callers: v2 wrongly read 2 of 98 caller turns as another voice; v1 read 0 of 98.
- A/B on 1,168 real turns:
- Caller turns flagged: v2 flagged 0/202 labelled PSTN turns and 3/487 web turns (likely handovers). v1 flagged 1/202 and 0/487.
- Real non-caller lines: there were only 6, and each model caught 1.
Same-gender speakers close to the phone remain the weak case. At 8 kHz, the ECAPA enrollment embedding itself separates same-gender voices poorly (EER 4β9 % on public speakers and about 16 % on real calls at 1 s).
Calibration
pvad_stream_v2.onnx has an offset of β1.5 on the "other" logit baked into
out.bias[2]. The offset is the least negative one that keeps, on validation:
- synthetic caller-utterance drops β€ 0.5 %;
- caller frames read as "other" β€ 3 %;
- real held-out-caller drops β€ 0.5 %.
pvad_v2.pt is the uncalibrated checkpoint (step 34,000). Add β1.5 to
out.bias[2] to reproduce the ONNX. The validation numbers are in
calibration.json.
Files
| file | what |
|---|---|
pvad_stream_v2.onnx |
streaming graph: one 10 ms frame per call, carries its own state, calibrated (use this in production) |
pvad_v2.pt |
PyTorch state_dict (uncalibrated; includes mu/sd) |
pvad_norm.json |
feature mean / sd (baked into the ONNX; here for retraining) |
calibration.json |
chosen step, the baked offset and its validation numbers |
config.json |
the full I/O, feature and evaluation contract |
example_stream.py |
reference streaming inference with onnxruntime |
train_log.txt |
the training run's loss and validation curve |
Model
fbank(40) β Linear(40β64) + ReLU
β causal Conv1d(64, k=5) + ReLU
β causal Conv1d(64, k=5, dil=2) + ReLU
β FiLM(emb 192 β scale, shift) (x Β· (1+Ξ³) + Ξ²)
β GRU(64β96, 2 layers)
β Linear(96β3) [silence, caller, other]
Input features
The input is a Kaldi-style log-mel filterbank, identical to v1:
- 8 kHz, 40 bins;
- 25 ms Hamming window, 10 ms hop, 256-point FFT;
- no dither, DC removal, 0.97 pre-emphasis;
lnenergies, no mean subtraction;- computed on int16-scale samples.
The exact settings are in config.json and example_stream.py.
Conditioning (enrollment) embedding
The enrollment is a 192-d, L2-normalised WeSpeaker ECAPA-TDNN512-LM
embedding (voxceleb_ECAPA512_LM.onnx, CC-BY-4.0) of 1.5β3 s of the caller's
clean speech: 8 β 16 kHz, 80-bin Kaldi fbank, per-utterance mean subtraction.
Any other speaker encoder requires retraining.
Streaming contract (pvad_stream_v2.onnx)
| tensor | shape | meaning |
|---|---|---|
feats |
[1,1,40] |
the new 10 ms frame |
hctx |
[1,4,64] |
last 4 post-ReLU input-hidden frames (zero-init) |
cctx |
[1,8,64] |
last 8 conv-1 outputs (zero-init) |
emb |
[1,192] |
enrollment embedding |
h0 |
[2,1,96] |
GRU state (zero-init) |
β logits |
[1,1,3] |
silence / caller / other |
β hn |
[2,1,96] |
next GRU state |
β h1 |
[1,1,64] |
push into hctx (drop the oldest) |
β c1 |
[1,1,64] |
push into cctx (drop the oldest) |
Reset the three states to zero at the start of each speech run after a pause longer than about 1 s. Training clips always start from the zero state.
Intended use and limits
- Purpose. v2 is built for barge-in gating in a full-duplex voice agent on telephony audio. It says who is speaking, not whom they address.
- Gate policy. Put a policy on top of the per-frame labels: for example, admit a non-caller voice after N speech runs within a window, and give way whenever a separate speaker check says the caller is talking.
- Wideband audio. Wideband and non-telephony audio is out of distribution. Downsample to 8 kHz first.
- Enrollment. The model needs an enrollment. Pass audio through until the caller has spoken for about 1.5 s over two turns.
- Prefer v1 where caller safety matters most. v2 is stronger on synthetic bystanders, TV and echo. On real calls it has not yet shown a gain over v1, and it drops slightly more caller turns. Run it in shadow, or behind a kill switch, before relying on it.
- Downloads last month
- 20