Nocturne v1 (Mini) β€” Distilled EfficientNet-B1 for Edge Bioacoustic ID

Nocturne Mini is the distilled, CPU-runnable student of stratus-labs/nocturne-v1-teacher. Same 2182-species vocab, ~9x smaller (9.3M vs 86M params), same ranking quality (mAP 0.163 vs teacher's 0.152 β€” student edges the teacher on mAP despite the size cut).

Use this variant when:

  • You need CPU / Apple Silicon inference at real-time budgets.
  • You're deploying to edge devices (Raspberry Pi 5, mobile, embedded audio recorders).
  • You want to run continuous acoustic monitoring on a low-power box.

Run it

The hosted demo and API were retired on 2026-09-27. This 9.4M-parameter checkpoint is built to run on a CPU or edge device: see Inference (Python) below.

How the three checkpoints compare

All three are scored by the same frozen harness on the same 13,710 test rows and the same 2,182-column class set, so the test columns are directly comparable between models.

checkpoint params test macro-F1 (calibrated) test mAP unseen-site anuran top-1 latency / 10 s clip (M4 Max)
v1 (teacher) 86M 0.1488 0.1498 0.544 225 ms
mini (this model) 9.4M 0.1542 0.1608 0.313 21 ms
v1.2 (teacher) 86M 0.1511 0.1541 0.138 225 ms

Two things worth stating plainly. First, the 9.4M student scores higher than either 86M teacher on the frozen test, at roughly a ninth of the parameters and about ten times the speed. Second, distillation appears to preserve generality: the student keeps far more unseen-site skill than the v1.2 teacher (0.313 vs 0.138), though still less than v1 itself (0.544). The unseen-site column is a diagnostic, not a gate β€” it is a much smaller slice (2,209 rows, 5 scorable classes) and is not row-comparable between models. If you are deploying to a brand-new recording site with no local validation, v1 is still the safer pick.

Model at a glance

Backbone EfficientNet-B1 (via timm), 1-channel log-mel input
Head Linear over 2182 species (multi-label BCE)
Input 10-second mono waveform @ 16 kHz β†’ 128-band log-mel image
Params ~9.3M
Precision bf16 for training; fp32 or bf16 for inference (portable)
Distilled from stratus-labs/nocturne-v1-teacher (AST, 86M)
License (weights) CC-BY-4.0
License (code) Apache-2.0

Evaluation

Real numbers from held-out val + test. These are the corrected numbers at threshold 0.3 (default) and with per-class thresholds calibrated on val (calibrated thresholds shipped as thresholds.json in this repo).

split metric threshold 0.3 calibrated per-class
val (13,391) macro-F1 0.129 0.164
micro-F1 0.485 0.526
mAP 0.163 0.163
test (13,710) macro-F1 0.121 0.154
micro-F1 0.490 0.512
mAP 0.161 0.161

Head-to-head with the AST teacher on the same test split:

model params test macro-F1 (calibrated) test mAP
nocturne-v1-teacher (AST) 86M 0.149 0.150
nocturne-v1-mini (this, EffNet-B1) 9.3M 0.154 0.161

Yes, the 9.3M-param distilled student edges the 86M-param teacher β€” soft-target training + threshold calibration line up nicely on this task. Use the mini for ~everything unless you specifically want the teacher's higher confidence at threshold 0.5+.

Reproduce with:

python -m soundscape.calibrate_and_eval \
  --config soundscape/configs/ast_nonbird.yaml \
  --checkpoint model.safetensors --arch effnet \
  --out-dir report/

vs BirdNET on non-bird taxa

Head-to-head on 300 random non-bird test clips. Top-1 species identification:

model non-bird top-1 accuracy
BirdNET (v2.4) 6.7%
Nocturne v1.1 teacher 76.7%
Nocturne v1 mini (this) β‰ˆ teacher (mini matches teacher on mAP; expect similar top-1)

11.5Γ— lift over BirdNET on the non-bird half of the soundscape. This is what the model is for.

Training recipe (distillation)

  • Loss: 0.2 * KL(sigmoid(teacher/T=1), sigmoid(student/T=1)) + 0.8 * BCE(student, hard_labels) β€” hard-label CE dominates; temperature=1 (no distribution flattening on multi-label Bernoulli targets).
  • Optimizer: AdamW lr 1e-3, cosine schedule, weight decay 0.01.
  • Augmentation: SpecAugment, MixUp (Ξ±=0.2).
  • Sampler: √-frequency class-balanced (same as teacher).
  • 30 epochs, best-of-N by val macro-F1.
  • Trained on 1Γ— NVIDIA GB10 (DGX Spark), bf16 mixed precision.

Inference (Python)

model.py in this repo is a standalone loader (torch, torchaudio, timm, safetensors, soundfile). Strict load; CPU is fine.

from huggingface_hub import snapshot_download
import sys

path = snapshot_download("stratus-labs/nocturne-v1-mini")
sys.path.insert(0, path)
from model import load_nocturne_mini, predict_file

model, vocab, thresholds = load_nocturne_mini(path)     # raises on any missing/unexpected key
for species, score in predict_file(model, "clip.wav", vocab, top_k=5):
    print(f"{score:.3f}  {species}")

Everything else

Coverage, intended use, out-of-scope caveats, ethical considerations, and citation are identical to the teacher card β€” see stratus-labs/nocturne-v1-teacher for the full text.

Research page

The write-up and every Nocturne release, in one place: https://runstratus.com/research/nocturne-v1-non-bird-bioacoustics

Downloads last month
58
Safetensors
Model size
9.4M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support