You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

pyannote segmentation 3.0

pyannote.audio's speaker segmentation model (SincNet + BiLSTM, powerset output), exported for loom.cpp. Family 13: audio in, for every ~17 ms frame a distribution over which of up to three speakers are talking.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from pyannote/segmentation-3.0. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

mit, inherited from the base model above.

Language(s)

(none tagged upstream)

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/pyannote-segmentation-3.0-loom")

# Audio is a mono float list at 16 kHz. The model was trained on 10-second windows; a longer
# recording is processed by sliding one over it.
window = audio[:10 * 16000]
result = model.speech2class.infer(window)
print(result.labels)
# ['non_speech', 'speaker1', 'speaker2', 'speaker3', 'speaker1+speaker2', 'speaker1+speaker3', 'speaker2+speaker3']
print(len(result), "frames,", round(result.frame_rate, 2), "per second")

# Each frame's most likely speaker SET, collapsed into turns. The speaker numbers are local to this
# window: telling who is who across windows takes an embedding model and a clustering step.
turns = []
for t, label in zip(result.times, result.best):
    if not turns or turns[-1][1] != label:
        turns.append((round(t, 2), label))
for start, label in turns:
    print(f"{start:6.2f}s  {label}")

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

Trained on 10-second windows. The export accepts other lengths, but what the model learned is 10 s of audio at a time, and a full recording is processed by sliding that window over it -- a loop the host owns, as pyannote.audio's own pipeline does.

Speaker labels are LOCAL to a window. speaker1 in one window and speaker1 in the next are not known to be the same person; this model segments, it does not identify. Full diarization adds a speaker embedding per segment (see titanet-large-loom) and a clustering step across windows, as the upstream card explains.

The seven classes are speaker SETS (pyannote's powerset encoding): no speech, one of three speakers alone, or one of the three pairs. Audio must be mono 16 kHz.

Upstream gates its repo behind an accept-terms form although the licence is MIT; this copy is gated the same way.

Files

  • pyannote-segmentation-3.0.gguf -- the model, exported with loom-exporter.
Downloads last month
1
GGUF
Model size
1.49M params
Architecture
loom-pyannote-segmentation
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/pyannote-segmentation-3.0-loom

Quantized
(10)
this model

Collection including loom-ai-org/pyannote-segmentation-3.0-loom