Philoria call recogniser: Perch 2 + random forest (V1)

Detects calls of Philoria in field recordings. Each one-second chunk of audio is turned into a Perch 2 embedding (Google DeepMind's model for animal sounds), and a random forest scores the embedding for the target call. Chunks scoring above 0.312 count as detections.

A fine-tuned Whisper recogniser for the same species is at LBolitho/Philoria_Call_Recogniser_V1. The Multi-Species Call Recogniser notebook can run either model, or both.

Files

File What it is
perch_rf.joblib the trained random forest (scikit-learn)
perch_rf_info.json the settings the model must be used with

How the audio must be prepared

  • Convert to mono at 32 kHz.
  • Cut into 1000 ms clips; centre each clip in a zero-padded 5 s window.
  • Embed with Perch 2 (perch_hoplite.zoo.model_configs.load_model_by_name("perch_v2")) and average the embedding over frames.
  • Score with rf.predict_proba(embeddings)[:, 1].

Training

  • Random forest: 500 trees
  • Classes: {'0': 'Non_Target_sounds', '1': 'Target_sounds'}
  • scikit-learn 1.6.1 (load the model with this version)

Validation

On held-out one-second clips from the training recordings:

| threshold | precision | recall | f1 | npv | roc_auc | avg_precision | |--------------------- | 0.312 | 0.990 | 0.990 | 0.990 | 0.999 | 0.999 | 0.997 | | 0.500 | 1.000 | 0.939 | 0.968 | 0.991 | 0.999 | 0.997 |

With whole recordings held out (grouped 5-fold cross-validation): F1 0.962, average precision 0.998.

Limitations

Tested only in the habitats where the training recordings were made. Validation clips come from the same recordings as the training clips, so field performance should be checked with manual review.

perch_rf.joblib is a pickle file: only load it from this repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support