Instructions to use stratus-labs/nocturne-v1-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use stratus-labs/nocturne-v1-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="stratus-labs/nocturne-v1-mini")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("stratus-labs/nocturne-v1-mini", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Nocturne v1 (Mini) β Distilled EfficientNet-B1 for Edge Bioacoustic ID
Nocturne Mini is the distilled, CPU-runnable student of stratus-labs/nocturne-v1-teacher. Same 2182-species vocab, ~9x smaller (9.3M vs 86M params), same ranking quality (mAP 0.163 vs teacher's 0.152 β student edges the teacher on mAP despite the size cut).
Use this variant when:
- You need CPU / Apple Silicon inference at real-time budgets.
- You're deploying to edge devices (Raspberry Pi 5, mobile, embedded audio recorders).
- You want to run continuous acoustic monitoring on a low-power box.
Run it
The hosted demo and API were retired on 2026-09-27. This 9.4M-parameter checkpoint is built to run on a CPU or edge device: see Inference (Python) below.
How the three checkpoints compare
All three are scored by the same frozen harness on the same 13,710 test rows and the same 2,182-column class set, so the test columns are directly comparable between models.
| checkpoint | params | test macro-F1 (calibrated) | test mAP | unseen-site anuran top-1 | latency / 10 s clip (M4 Max) |
|---|---|---|---|---|---|
v1 (teacher) |
86M | 0.1488 | 0.1498 | 0.544 | 225 ms |
mini (this model) |
9.4M | 0.1542 | 0.1608 | 0.313 | 21 ms |
v1.2 (teacher) |
86M | 0.1511 | 0.1541 | 0.138 | 225 ms |
Two things worth stating plainly. First, the 9.4M student scores higher than either 86M teacher on the
frozen test, at roughly a ninth of the parameters and about ten times the speed. Second, distillation
appears to preserve generality: the student keeps far more unseen-site skill than the v1.2 teacher (0.313
vs 0.138), though still less than v1 itself (0.544). The unseen-site column is a diagnostic, not a gate
β it is a much smaller slice (2,209 rows, 5 scorable classes) and is not row-comparable between models.
If you are deploying to a brand-new recording site with no local validation, v1 is still the safer pick.
Model at a glance
| Backbone | EfficientNet-B1 (via timm), 1-channel log-mel input |
| Head | Linear over 2182 species (multi-label BCE) |
| Input | 10-second mono waveform @ 16 kHz β 128-band log-mel image |
| Params | ~9.3M |
| Precision | bf16 for training; fp32 or bf16 for inference (portable) |
| Distilled from | stratus-labs/nocturne-v1-teacher (AST, 86M) |
| License (weights) | CC-BY-4.0 |
| License (code) | Apache-2.0 |
Evaluation
Real numbers from held-out val + test. These are the corrected numbers at threshold 0.3 (default) and with per-class thresholds calibrated on val (calibrated thresholds shipped as thresholds.json in this repo).
| split | metric | threshold 0.3 | calibrated per-class |
|---|---|---|---|
| val (13,391) | macro-F1 | 0.129 | 0.164 |
| micro-F1 | 0.485 | 0.526 | |
| mAP | 0.163 | 0.163 | |
| test (13,710) | macro-F1 | 0.121 | 0.154 |
| micro-F1 | 0.490 | 0.512 | |
| mAP | 0.161 | 0.161 |
Head-to-head with the AST teacher on the same test split:
| model | params | test macro-F1 (calibrated) | test mAP |
|---|---|---|---|
nocturne-v1-teacher (AST) |
86M | 0.149 | 0.150 |
nocturne-v1-mini (this, EffNet-B1) |
9.3M | 0.154 | 0.161 |
Yes, the 9.3M-param distilled student edges the 86M-param teacher β soft-target training + threshold calibration line up nicely on this task. Use the mini for ~everything unless you specifically want the teacher's higher confidence at threshold 0.5+.
Reproduce with:
python -m soundscape.calibrate_and_eval \
--config soundscape/configs/ast_nonbird.yaml \
--checkpoint model.safetensors --arch effnet \
--out-dir report/
vs BirdNET on non-bird taxa
Head-to-head on 300 random non-bird test clips. Top-1 species identification:
| model | non-bird top-1 accuracy |
|---|---|
| BirdNET (v2.4) | 6.7% |
| Nocturne v1.1 teacher | 76.7% |
| Nocturne v1 mini (this) | β teacher (mini matches teacher on mAP; expect similar top-1) |
11.5Γ lift over BirdNET on the non-bird half of the soundscape. This is what the model is for.
Training recipe (distillation)
- Loss:
0.2 * KL(sigmoid(teacher/T=1), sigmoid(student/T=1)) + 0.8 * BCE(student, hard_labels)β hard-label CE dominates; temperature=1 (no distribution flattening on multi-label Bernoulli targets). - Optimizer: AdamW lr 1e-3, cosine schedule, weight decay 0.01.
- Augmentation: SpecAugment, MixUp (Ξ±=0.2).
- Sampler: β-frequency class-balanced (same as teacher).
- 30 epochs, best-of-N by val macro-F1.
- Trained on 1Γ NVIDIA GB10 (DGX Spark), bf16 mixed precision.
Inference (Python)
model.py in this repo is a standalone loader (torch, torchaudio, timm, safetensors, soundfile). Strict load; CPU is fine.
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("stratus-labs/nocturne-v1-mini")
sys.path.insert(0, path)
from model import load_nocturne_mini, predict_file
model, vocab, thresholds = load_nocturne_mini(path) # raises on any missing/unexpected key
for species, score in predict_file(model, "clip.wav", vocab, top_k=5):
print(f"{score:.3f} {species}")
Everything else
Coverage, intended use, out-of-scope caveats, ethical considerations, and citation are identical to the teacher card β see stratus-labs/nocturne-v1-teacher for the full text.
Research page
The write-up and every Nocturne release, in one place: https://runstratus.com/research/nocturne-v1-non-bird-bioacoustics
- Downloads last month
- 58