mosquito-mtrcnn-dg

Mosquito logo: a mosquito seen from the side on a green tile

A compact MTRCNN classifier that names the mosquito species in a flight-tone recording, trained for cross-domain generalization on the BioDCASE 2026 Task 5 development data. It knows 9 species.

Research aid only. Not for public-health or vector-control decisions. Accuracy on recordings from unseen setups is low; see Results.

🚀 Usage

hf download jgalego/mosquito-mtrcnn-dg mtrcnn.py --local-dir .
uv run mtrcnn.py predict clip.wav

predict downloads model.pt, resamples each WAV file to 8 kHz, splits it into 2-second windows, averages their embeddings, and prints the species whose class mean is nearest in Mahalanobis distance. It takes any number of files. mtrcnn.py also holds the training and evaluation code.

Species: Aedes aegypti, Aedes albopictus, Culex quinquefasciatus, Anopheles gambiae, Anopheles arabiensis, Anopheles dirus, Culex pipiens, Anopheles minimus, Anopheles stephensi.

🏋️ Training

Data All 213,647 clips of the official training split of BioDCASE 2026 Challenge: Cross-Domain Mosquito Species Classification (Hou et al., Zenodo, 10.5281/zenodo.20478577, CC BY 4.0)
Model MTRCNN with three temporal-resolution branches and a 32-dimensional embedding
Domain generalization FourierMix and MixStyle across recording domains, a gradient-reversal domain classifier, class-conditional domain losses, supervised contrastive loss
Input 2-second crops, 64-bin log-mel up to 4 kHz
Sampling Balanced over species-domain cells, logit adjustment with the sampler's species priors
Schedule 12 epochs, AdamW, learning rate 0.001, batch size 128, cosine decay; the last epoch is kept
Inference Mahalanobis distance to class means of 20,000 training clips with one shared covariance matrix
Seed 42

📊 Results

Balanced accuracy (%) on all 27,217 clips of the official development test split. Unseen clips (4,202) come from species-domain pairs that do not occur in training.

This model Seen Unseen
Mahalanobis 74.90 28.25
Softmax 69.32 25.47

Mean ± standard deviation over training seeds:

Configuration Inference Seeds Seen Unseen
2-second crops Softmax 3 67.04 ± 2.40 25.15 ± 0.53
2-second crops Mahalanobis 3 73.76 ± 1.14 27.74 ± 0.45
0.625-second crops, RMS-normalized, mel from 200 Hz Softmax 3 76.77 ± 0.23 20.57 ± 0.50
0.625-second crops, RMS-normalized, mel from 200 Hz Mahalanobis 3 78.09 ± 0.33 20.72 ± 0.69
2-second crops, liftered log-mel Softmax 3 74.60 ± 0.45 21.70 ± 0.49
2-second crops, liftered log-mel Mahalanobis 3 75.78 ± 0.96 21.76 ± 1.03

For reference, on the same split:

System Seen Unseen
Official MTRCNN baseline, 10 seeds (task page) 88.06 ± 1.08 17.51 ± 1.97
Fine-tuned WavLM-Large 23.74 17.25
Probes on frozen Perch v2 embeddings, per seed † 31-35
Our probes on Perch v2 embeddings, 8 seeds 82.33 28.78
Our probes on harmonic features, 8 seeds 72.49 19.08
Our probes on background-whitened features, 8 seeds 74.21 18.34
Equal-weight ensemble of the three probes and this model 81.17 24.41
Agreement-gated Perch, harmonic, BirdMAE and background-whitened ensemble † 36.16

† Reported in aptemvs/mosquitoes-biodcase2026-task5 (legacy/final/) and not reproduced here. The checkpoints only accept precomputed Perch v2 and BirdMAE embeddings, the inference script depends on code that is not in that repository, and the ensemble's design was compared against this same test split, so these figures are likely optimistic.

Our probes use the precomputed features from the same dataset, with domain-balanced sampling and domain alignment losses. They did not reproduce the † figures, and averaging them with this model did not beat it, so we kept the model as is. Aedes aegypti stays at 0% for every system.

This model generalizes better than the MTRCNN baseline but gives up about 14 points of seen-domain accuracy.

Unseen recall per species, Mahalanobis:

Species Recall
Aedes aegypti 0%
Aedes albopictus 59%
Culex quinquefasciatus 5%
Anopheles gambiae 0%
Anopheles arabiensis 0%
Anopheles dirus 0%
Culex pipiens 94%
Anopheles minimus 74%
Anopheles stephensi 22%

⚠️ Limitations

  • 99.4% of the training audio comes from one recording domain. Anopheles gambiae and Anopheles arabiensis have no training clips from any other domain.
  • On unseen domains the model tends to predict a species that the training data has in the same domain, so several species are almost never recognized there.
  • The configurations above were compared on this development test split, so the numbers are slightly optimistic for new recordings.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train jgalego/mosquito-mtrcnn-dg

Space using jgalego/mosquito-mtrcnn-dg 1

Collection including jgalego/mosquito-mtrcnn-dg

Evaluation results

  • Unseen-domain balanced accuracy on BioDCASE 2026 Task 5 development test, unseen domains
    self-reported
    28.250