pVAD v2: speaker-conditioned frame VAD for telephony barge-in

pVAD v2 labels every 10 ms of an 8 kHz caller stream as silence, the enrolled caller, or someone else. "Someone else" covers a bystander, a TV or laptop, or the agent's own echo. The model is conditioned on a 192-d speaker embedding of the caller.

v2 uses the same 171k-parameter graph and streaming contract as v1, so it is a drop-in replacement. The difference is the training data:

  • v1 learned mostly near/far loudness.
  • v2 was trained on 90k synthetic scenes built to separate who is speaking:
    • the speakers: same- and other-gender bystanders, TV and laptop playback, agent echo, the caller on speakerphone, far rooms, music, babble, short words and handovers;
    • the channels: G.711, AMR, Opus, G.722 and web audio;
    • the voices: public multilingual speech plus caller voices from Telenow's own Indian call recordings.

No audio is released with this model.

Results

Synthetic (held-out speakers)

Each cell has 250 six-second mixtures of speakers unseen in training. Each model runs at its own caller-safe calibration (see below). The table gives the share of another voice's utterances that the model flags as not the caller.

scenario v1 v2
agent echo (caller silent / double-talk) 1.0 % 56.7 %
near-field person, other gender 0 % 48.6 %
bystander/TV only, and handovers 0.9 % 44.2 %
TV 1–3 m 0 % 40.5 %
laptop/phone 0.2–0.8 m 0 % 36.8 %
caller on speakerphone, bystander close 0 % 25.6 %
far room 1–4 m 2.5 % 23.9 %
near-field person, same gender 0 % 7.5 %
same gender, low pitch 130–180 Hz 0 % 8.8 %
enrollment 1.5–2 s, 20–40 % contaminated 0 % 4.0 %

Caller utterances dropped: at most 1.2 % in any cell (caller-only cell: 0.4 %).

Real calls

On real calls v2 has not yet proven better than v1:

  • Held-out callers: v2 wrongly read 2 of 98 caller turns as another voice; v1 read 0 of 98.
  • A/B on 1,168 real turns:
    • Caller turns flagged: v2 flagged 0/202 labelled PSTN turns and 3/487 web turns (likely handovers). v1 flagged 1/202 and 0/487.
    • Real non-caller lines: there were only 6, and each model caught 1.

Same-gender speakers close to the phone remain the weak case. At 8 kHz, the ECAPA enrollment embedding itself separates same-gender voices poorly (EER 4–9 % on public speakers and about 16 % on real calls at 1 s).

Calibration

pvad_stream_v2.onnx has an offset of βˆ’1.5 on the "other" logit baked into out.bias[2]. The offset is the least negative one that keeps, on validation:

  • synthetic caller-utterance drops ≀ 0.5 %;
  • caller frames read as "other" ≀ 3 %;
  • real held-out-caller drops ≀ 0.5 %.

pvad_v2.pt is the uncalibrated checkpoint (step 34,000). Add βˆ’1.5 to out.bias[2] to reproduce the ONNX. The validation numbers are in calibration.json.

Files

file what
pvad_stream_v2.onnx streaming graph: one 10 ms frame per call, carries its own state, calibrated (use this in production)
pvad_v2.pt PyTorch state_dict (uncalibrated; includes mu/sd)
pvad_norm.json feature mean / sd (baked into the ONNX; here for retraining)
calibration.json chosen step, the baked offset and its validation numbers
config.json the full I/O, feature and evaluation contract
example_stream.py reference streaming inference with onnxruntime
train_log.txt the training run's loss and validation curve

Model

fbank(40) β†’ Linear(40β†’64) + ReLU
          β†’ causal Conv1d(64, k=5)           + ReLU
          β†’ causal Conv1d(64, k=5, dil=2)    + ReLU
          β†’ FiLM(emb 192 β†’ scale, shift)     (x Β· (1+Ξ³) + Ξ²)
          β†’ GRU(64β†’96, 2 layers)
          β†’ Linear(96β†’3)   [silence, caller, other]

Input features

The input is a Kaldi-style log-mel filterbank, identical to v1:

  • 8 kHz, 40 bins;
  • 25 ms Hamming window, 10 ms hop, 256-point FFT;
  • no dither, DC removal, 0.97 pre-emphasis;
  • ln energies, no mean subtraction;
  • computed on int16-scale samples.

The exact settings are in config.json and example_stream.py.

Conditioning (enrollment) embedding

The enrollment is a 192-d, L2-normalised WeSpeaker ECAPA-TDNN512-LM embedding (voxceleb_ECAPA512_LM.onnx, CC-BY-4.0) of 1.5–3 s of the caller's clean speech: 8 β†’ 16 kHz, 80-bin Kaldi fbank, per-utterance mean subtraction. Any other speaker encoder requires retraining.

Streaming contract (pvad_stream_v2.onnx)

tensor shape meaning
feats [1,1,40] the new 10 ms frame
hctx [1,4,64] last 4 post-ReLU input-hidden frames (zero-init)
cctx [1,8,64] last 8 conv-1 outputs (zero-init)
emb [1,192] enrollment embedding
h0 [2,1,96] GRU state (zero-init)
β†’ logits [1,1,3] silence / caller / other
β†’ hn [2,1,96] next GRU state
β†’ h1 [1,1,64] push into hctx (drop the oldest)
β†’ c1 [1,1,64] push into cctx (drop the oldest)

Reset the three states to zero at the start of each speech run after a pause longer than about 1 s. Training clips always start from the zero state.

Intended use and limits

  • Purpose. v2 is built for barge-in gating in a full-duplex voice agent on telephony audio. It says who is speaking, not whom they address.
  • Gate policy. Put a policy on top of the per-frame labels: for example, admit a non-caller voice after N speech runs within a window, and give way whenever a separate speaker check says the caller is talking.
  • Wideband audio. Wideband and non-telephony audio is out of distribution. Downsample to 8 kHz first.
  • Enrollment. The model needs an enrollment. Pass audio through until the caller has spoken for about 1.5 s over two turns.
  • Prefer v1 where caller safety matters most. v2 is stronger on synthetic bystanders, TV and echo. On real calls it has not yet shown a gain over v1, and it drops slightly more caller turns. Run it in shadow, or behind a kill switch, before relying on it.
Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support