WhisperSeg AER checkpoints
WhisperSeg is highly effective at segmenting animal vocalizations, but its original checkpoints are large and autoregressive decoding can be slow. WhisperSeg AER makes these models accessible to computers with less GPU memory through two steps:
- Compact the original models without retraining. Keep the numeric and special tokens needed for segmentation, remove unused vocabulary rows, convert weights to FP16, and package the model and tokenizer in one DAS-compatible
.ckptfile. The encoder and decoder depths stay the same. On the tests below, the compact and original models perform virtually identically. - Prune and retrain the Base decoder. Starting from compact Animal VAD Base, keep all six encoder layers but reduce the decoder from six layers to four, two, or one. Train the remaining decoder on the
vad-animalstraining split with the encoder frozen. Pruning makes the Base checkpoints smaller and inference faster, but may limit transfer to unfamiliar sounds. Four and two decoder layers retain nearly the same multispecies F1 as the retrained six-layer model here; nightingale transfer varies more, particularly with one layer.
All seven release checkpoints are packaged for DAS 1.0a1.
| Model | Checkpoint |
|---|---|
| Large multispecies (compact FP16) | whisperseg-large-ms-cmp-fp16.ckpt |
| Animal VAD Large (compact FP16) | whisperseg-animal-vad-cmp-fp16.ckpt |
| Animal VAD Base (compact FP16) | whisperseg-base-animal-vad-cmp-fp16.ckpt |
| Base 6E+6D, retrained | whisperseg-base-animal-vad-6e6d-cmp-fp16.ckpt |
| Base 6E+4D | whisperseg-base-animal-vad-6e4d-cmp-fp16.ckpt |
| Base 6E+2D | whisperseg-base-animal-vad-6e2d-cmp-fp16.ckpt |
| Base 6E+1D | whisperseg-base-animal-vad-6e1d-cmp-fp16.ckpt |
cmp means compact vocabulary; all seven files contain FP16 weights. The 6E+6D filename distinguishes the decoder-retrained Base model from the non-retrained compact Base.
Use with DAS 1.0a1
Install DAS 1.0a1, download a checkpoint, then run inference on one second of zero-valued audio:
import numpy as np
from das.whisperseg.model import WhisperSegmenter
segmenter = WhisperSegmenter("whisperseg-base-animal-vad-cmp-fp16.ckpt", device="cpu")
prediction = segmenter.segment(
np.zeros(16_000, dtype=np.float32),
sr=16_000,
num_trials=1,
batch_size=1,
num_beams=1,
max_length=128,
)
Evaluation
The three original FP32 WhisperSeg models are comparison references, not files in this repository. All ten models used max_length=128 on the same files. Multispecies F1 averages four species equally on a fixed 65-file, 920-event subset of the vad-animals test split; nightingale F1 uses one 945-second, 2,870-event recording that was not used for training. Both use one-to-one event matching at IoU ≥0.5. Memory and speed were measured during nightingale inference on an Apple M2 Max using MPS, batch size 1, three trials, and four beams. Each model and dataset ran in a fresh process after an idle-host check. Release checkpoints use FP16; originals use FP32. The original model sources are Animal VAD Base, Animal VAD Large, and Large multispecies.
| Model | Derived from | Decoder retrained here | Encoder / decoder | Parameters (M) | File (MiB) | Multispecies macro F1 ↑ | Nightingale F1 ↑ | Night peak MPS (GiB) ↓ | Night inference s/audio s ↓ |
|---|---|---|---|---|---|---|---|---|---|
| Animal VAD Base · original reference | Original Animal VAD Base | No | 6 / 6 | 72.1 | 275.1 | 0.7041 | 0.5948 | 1.311 | 0.372 |
| Animal VAD Base · compact FP16 | Original Animal VAD Base | No | 6 / 6 | 46.1 | 89.5 | 0.7041 | 0.5950 | 0.277 | 0.335 |
| Base 6E+6D · retrained | Compact Animal VAD Base | Yes | 6 / 6 | 46.1 | 88.3 | 0.9721 | 0.6967 | 0.230 | 0.325 |
| Base 6E+4D | Compact Animal VAD Base | Yes | 6 / 4 | 37.7 | 72.3 | 0.9713 | 0.7177 | 0.168 | 0.268 |
| Base 6E+2D | Compact Animal VAD Base | Yes | 6 / 2 | 29.3 | 56.2 | 0.9708 | 0.6842 | 0.129 | 0.208 |
| Base 6E+1D | Compact Animal VAD Base | Yes | 6 / 1 | 25.1 | 48.2 | 0.9621 | 0.6597 | 0.129 | 0.184 |
| Animal VAD Large · original reference | Original Animal VAD Large | No | 32 / 32 | 1,542.0 | 5,882.8 | 0.7004 | 0.7413 | 10.297 | 2.367 |
| Animal VAD Large · compact FP16 | Original Animal VAD Large | No | 32 / 32 | 1,477.1 | 2,820.9 | 0.7015 | 0.7409 | 6.360 | 1.822 |
| Large multispecies · original reference | Original Large multispecies | No | 32 / 32 | 1,542.0 | 5,882.8 | 0.6997 | 0.5456 | 10.484 | 2.182 |
| Large multispecies · compact FP16 | Original Large multispecies | No | 32 / 32 | 1,477.1 | 2,820.9 | 0.7007 | 0.5445 | 6.555 | 1.680 |
Bold marks results within 0.1% of the best value in each column, using the arrows to indicate direction. Inerence run with max_length=128. MPS peak includes allocator and framework allocations and is not a guaranteed GPU-memory requirement on other hardware. Timing excludes model loading, audio-file reading, and scoring. Original file sizes refer to pytorch_model.bin; release sizes refer to .ckpt files.
Base 6E+4D is a useful starting point when both memory and transfer performance matter: 0.9713 multispecies F1 and 0.7177 nightingale F1 with 0.168 GiB observed MPS peak. Base 6E+1D is the smallest and fastest measured release model. Compact Animal VAD Large has the highest nightingale F1 among release checkpoints (0.7409), at substantially higher memory and runtime.
Citation
Please cite both papers when using any model in this repository:
- Gu et al. (2024), Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection, ICASSP 2024 (WhisperSeg).
- Steinfath et al. (2021), Fast and accurate annotation of acoustic signals with deep neural networks, eLife 10:e68837 (DAS).
Model tree for DAS-DeepAudioSegmenter/whisperseg-aer
Base model
nccratliri/whisperseg-animal-vad