WhisperSeg AER checkpoints

WhisperSeg is highly effective at segmenting animal vocalizations, but its original checkpoints are large and autoregressive decoding can be slow. WhisperSeg AER makes these models accessible to computers with less GPU memory through two steps:

  1. Compact the original models without retraining. Keep the numeric and special tokens needed for segmentation, remove unused vocabulary rows, convert weights to FP16, and package the model and tokenizer in one DAS-compatible .ckpt file. The encoder and decoder depths stay the same. On the tests below, the compact and original models perform virtually identically.
  2. Prune and retrain the Base decoder. Starting from compact Animal VAD Base, keep all six encoder layers but reduce the decoder from six layers to four, two, or one. Train the remaining decoder on the vad-animals training split with the encoder frozen. Pruning makes the Base checkpoints smaller and inference faster, but may limit transfer to unfamiliar sounds. Four and two decoder layers retain nearly the same multispecies F1 as the retrained six-layer model here; nightingale transfer varies more, particularly with one layer.

All seven release checkpoints are packaged for DAS 1.0a1.

Model Checkpoint
Large multispecies (compact FP16) whisperseg-large-ms-cmp-fp16.ckpt
Animal VAD Large (compact FP16) whisperseg-animal-vad-cmp-fp16.ckpt
Animal VAD Base (compact FP16) whisperseg-base-animal-vad-cmp-fp16.ckpt
Base 6E+6D, retrained whisperseg-base-animal-vad-6e6d-cmp-fp16.ckpt
Base 6E+4D whisperseg-base-animal-vad-6e4d-cmp-fp16.ckpt
Base 6E+2D whisperseg-base-animal-vad-6e2d-cmp-fp16.ckpt
Base 6E+1D whisperseg-base-animal-vad-6e1d-cmp-fp16.ckpt

cmp means compact vocabulary; all seven files contain FP16 weights. The 6E+6D filename distinguishes the decoder-retrained Base model from the non-retrained compact Base.

Use with DAS 1.0a1

Install DAS 1.0a1, download a checkpoint, then run inference on one second of zero-valued audio:

import numpy as np
from das.whisperseg.model import WhisperSegmenter

segmenter = WhisperSegmenter("whisperseg-base-animal-vad-cmp-fp16.ckpt", device="cpu")
prediction = segmenter.segment(
    np.zeros(16_000, dtype=np.float32),
    sr=16_000,
    num_trials=1,
    batch_size=1,
    num_beams=1,
    max_length=128,
)

Evaluation

The three original FP32 WhisperSeg models are comparison references, not files in this repository. All ten models used max_length=128 on the same files. Multispecies F1 averages four species equally on a fixed 65-file, 920-event subset of the vad-animals test split; nightingale F1 uses one 945-second, 2,870-event recording that was not used for training. Both use one-to-one event matching at IoU ≥0.5. Memory and speed were measured during nightingale inference on an Apple M2 Max using MPS, batch size 1, three trials, and four beams. Each model and dataset ran in a fresh process after an idle-host check. Release checkpoints use FP16; originals use FP32. The original model sources are Animal VAD Base, Animal VAD Large, and Large multispecies.

Model Derived from Decoder retrained here Encoder / decoder Parameters (M) File (MiB) Multispecies macro F1 ↑ Nightingale F1 ↑ Night peak MPS (GiB) ↓ Night inference s/audio s ↓
Animal VAD Base · original reference Original Animal VAD Base No 6 / 6 72.1 275.1 0.7041 0.5948 1.311 0.372
Animal VAD Base · compact FP16 Original Animal VAD Base No 6 / 6 46.1 89.5 0.7041 0.5950 0.277 0.335
Base 6E+6D · retrained Compact Animal VAD Base Yes 6 / 6 46.1 88.3 0.9721 0.6967 0.230 0.325
Base 6E+4D Compact Animal VAD Base Yes 6 / 4 37.7 72.3 0.9713 0.7177 0.168 0.268
Base 6E+2D Compact Animal VAD Base Yes 6 / 2 29.3 56.2 0.9708 0.6842 0.129 0.208
Base 6E+1D Compact Animal VAD Base Yes 6 / 1 25.1 48.2 0.9621 0.6597 0.129 0.184
Animal VAD Large · original reference Original Animal VAD Large No 32 / 32 1,542.0 5,882.8 0.7004 0.7413 10.297 2.367
Animal VAD Large · compact FP16 Original Animal VAD Large No 32 / 32 1,477.1 2,820.9 0.7015 0.7409 6.360 1.822
Large multispecies · original reference Original Large multispecies No 32 / 32 1,542.0 5,882.8 0.6997 0.5456 10.484 2.182
Large multispecies · compact FP16 Original Large multispecies No 32 / 32 1,477.1 2,820.9 0.7007 0.5445 6.555 1.680

Bold marks results within 0.1% of the best value in each column, using the arrows to indicate direction. Inerence run with max_length=128. MPS peak includes allocator and framework allocations and is not a guaranteed GPU-memory requirement on other hardware. Timing excludes model loading, audio-file reading, and scoring. Original file sizes refer to pytorch_model.bin; release sizes refer to .ckpt files.

Base 6E+4D is a useful starting point when both memory and transfer performance matter: 0.9713 multispecies F1 and 0.7177 nightingale F1 with 0.168 GiB observed MPS peak. Base 6E+1D is the smallest and fastest measured release model. Compact Animal VAD Large has the highest nightingale F1 among release checkpoints (0.7409), at substantially higher memory and runtime.

Citation

Please cite both papers when using any model in this repository:

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DAS-DeepAudioSegmenter/whisperseg-aer

Finetuned
(1)
this model

Dataset used to train DAS-DeepAudioSegmenter/whisperseg-aer