Silero VAD v6

Silero's voice activity detector (STFT + CNN + LSTM, ~300K parameters), exported for loom.cpp. Family 13: audio in, a speech probability for every 32 ms frame out.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from Silero VAD v6.2.3 (16 kHz). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

mit, inherited from the base model above.

Language(s)

(none tagged upstream)

Language-agnostic: upstream reports training on corpora covering over 6,000 languages.

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/silero-vad-loom")

# Audio is a mono float list at 16 kHz. One call over the whole clip, one row per 32 ms frame back.
result = model.speech2class.infer(audio)
print(result.granularity, result.labels)
# frame ['non_speech', 'speech']

# A probability per frame, not a decision: the threshold and any smoothing are yours. With a plain
# 0.5 threshold, the speech curve becomes segments:
speech = result.probability("speech")
segments, start = [], None
for t, p in zip(result.times, speech):
    if p >= 0.5 and start is None:
        start = t
    elif p < 0.5 and start is not None:
        segments.append((round(start, 2), round(t, 2)))
        start = None
if start is not None:
    segments.append((round(start, 2), round(len(audio) / 16000, 2)))
print(segments)   # [(start_seconds, end_seconds), ...]

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

The answer is a probability per frame, not a decision. Silero's own get_speech_timestamps thresholds it (0.5), pads each segment and drops short ones; that post-processing is left to you, because every consumer of a VAD wants different onset and offset behaviour.

One row per 512 samples (32 ms) of input, starting at 0 s: result.times gives each row's start. The whole clip is one call -- the model's streaming state is carried across every frame of it, exactly as upstream's frame loop carries it. Audio must be mono 16 kHz; upstream's 8 kHz model, in the same JIT, is not exported.

The weights are upstream's, RE-LAID OUT: the per-frame reflect pad is folded into the STFT kernel and the four per-frame convolutions become dense pointwise maps, so the whole clip is one pass. The arithmetic is the same -- checked against upstream's own JIT at every export -- and the probabilities agree with it to its own float rounding (1.6e-5 worst case on real speech, against an f64 reference it is itself 1.0e-5 from).

Files

  • silero-vad.gguf -- the model, exported with loom-exporter.
Downloads last month
16
GGUF
Model size
1.07M params
Architecture
loom-silero-vad
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including loom-ai-org/silero-vad-loom