Silero VAD v6
Silero's voice activity detector (STFT + CNN + LSTM, ~300K parameters), exported for loom.cpp. Family 13: audio in, a speech probability for every 32 ms frame out.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from Silero VAD v6.2.3 (16 kHz). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.
License
mit, inherited from the base model above.
Language(s)
(none tagged upstream)
Language-agnostic: upstream reports training on corpora covering over 6,000 languages.
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/silero-vad-loom")
# Audio is a mono float list at 16 kHz. One call over the whole clip, one row per 32 ms frame back.
result = model.speech2class.infer(audio)
print(result.granularity, result.labels)
# frame ['non_speech', 'speech']
# A probability per frame, not a decision: the threshold and any smoothing are yours. With a plain
# 0.5 threshold, the speech curve becomes segments:
speech = result.probability("speech")
segments, start = [], None
for t, p in zip(result.times, speech):
if p >= 0.5 and start is None:
start = t
elif p < 0.5 and start is not None:
segments.append((round(start, 2), round(t, 2)))
start = None
if start is not None:
segments.append((round(start, 2), round(len(audio) / 16000, 2)))
print(segments) # [(start_seconds, end_seconds), ...]
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
The answer is a probability per frame, not a decision. Silero's own get_speech_timestamps thresholds it (0.5), pads each segment and drops short ones; that post-processing is left to you, because every consumer of a VAD wants different onset and offset behaviour.
One row per 512 samples (32 ms) of input, starting at 0 s: result.times gives each row's start. The whole clip is one call -- the model's streaming state is carried across every frame of it, exactly as upstream's frame loop carries it. Audio must be mono 16 kHz; upstream's 8 kHz model, in the same JIT, is not exported.
The weights are upstream's, RE-LAID OUT: the per-frame reflect pad is folded into the STFT kernel and the four per-frame convolutions become dense pointwise maps, so the whole clip is one pass. The arithmetic is the same -- checked against upstream's own JIT at every export -- and the probabilities agree with it to its own float rounding (1.6e-5 worst case on real speech, against an f64 reference it is itself 1.0e-5 from).
Files
silero-vad.gguf-- the model, exported with loom-exporter.
- Downloads last month
- 16
We're not able to determine the quantization variants.