Update README.md
Browse files
README.md
CHANGED
|
@@ -1,4 +1,3 @@
|
|
| 1 |
-
|
| 2 |
---
|
| 3 |
license: apache-2.0
|
| 4 |
library_name: pytorch
|
|
@@ -11,72 +10,58 @@ tags:
|
|
| 11 |
- perch
|
| 12 |
---
|
| 13 |
|
| 14 |
-
#
|
| 15 |
|
| 16 |
-
|
| 17 |
|
|
|
|
| 18 |
|
| 19 |
## What this is
|
| 20 |
|
| 21 |
-
-
|
| 22 |
-
|
|
|
|
| 23 |
- **File:** `perch_v2_backbone_timm.pt` — backbone weights only (no classification head).
|
| 24 |
|
| 25 |
-
## Known limitation
|
|
|
|
|
|
|
| 26 |
|
| 27 |
-
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
-
|
| 32 |
|
| 33 |
-
##
|
|
|
|
|
|
|
| 34 |
|
| 35 |
```python
|
|
|
|
|
|
|
|
|
|
| 36 |
import torch
|
| 37 |
from huggingface_hub import hf_hub_download
|
| 38 |
-
from
|
| 39 |
|
| 40 |
weights_path = hf_hub_download(repo_id="bghani/perch2-pytorch-weights", filename="perch_v2_backbone_timm.pt")
|
| 41 |
|
|
|
|
| 42 |
embedder = Perch2Embedder(weights_path=weights_path)
|
| 43 |
embedder.eval()
|
| 44 |
-
|
| 45 |
waveform = torch.zeros(4, 160_000) # 5s clips @ 32kHz, batch of 4
|
| 46 |
with torch.no_grad():
|
| 47 |
embeddings = embedder(waveform) # (4, 1536)
|
| 48 |
-
```
|
| 49 |
-
|
| 50 |
-
### 2. Linear probing
|
| 51 |
-
|
| 52 |
-
```python
|
| 53 |
-
from perchv2_pytorch import PerchFrontend, Perch2Classifier
|
| 54 |
|
|
|
|
| 55 |
mel = PerchFrontend()
|
| 56 |
-
model = Perch2Classifier(
|
| 57 |
-
num_classes=42,
|
| 58 |
-
mel=mel,
|
| 59 |
-
weights_path=weights_path,
|
| 60 |
-
mode="linear_probe", # backbone frozen, only the new head trains
|
| 61 |
-
)
|
| 62 |
-
```
|
| 63 |
-
|
| 64 |
-
### 3. Full fine-tuning
|
| 65 |
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
mel = PerchFrontend()
|
| 70 |
-
model = Perch2Classifier(
|
| 71 |
-
num_classes=42,
|
| 72 |
-
mel=mel,
|
| 73 |
-
weights_path=weights_path,
|
| 74 |
-
mode="finetune", # backbone unfrozen -- this is the whole point of this repo
|
| 75 |
-
)
|
| 76 |
```
|
| 77 |
|
| 78 |
-
See the [main repo](https://github.com/bghani/perchv2-pytorch) for full usage — frozen embeddings, linear probing, and full fine-tuning, with runnable examples and a walkthrough notebook.
|
| 79 |
-
|
| 80 |
## License
|
| 81 |
|
| 82 |
Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://github.com/bghani/perchv2-pytorch/blob/main/NOTICE) in the main repo for the full derivative-work attribution.
|
|
@@ -84,25 +69,11 @@ Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://g
|
|
| 84 |
## Citation
|
| 85 |
|
| 86 |
If you use this in published work, please cite the original Perch v2 paper:
|
| 87 |
-
|
| 88 |
```
|
| 89 |
@article{van2025perch,
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
}
|
| 95 |
```
|
| 96 |
-
If you use this specific PyTorch conversion, please also cite [ChiroEcho](https://arxiv.org/abs/2608.18191), where it was first developed and applied (CV4Ecology - ECCV 2026):
|
| 97 |
-
|
| 98 |
-
```
|
| 99 |
-
@misc{ghani2026chiroechoextendingautomatedbat,
|
| 100 |
-
title={ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy},
|
| 101 |
-
author={Burooj Ghani and Welmoed Eversteijn and Milan van Hirtum and Juan Sebastián Cañas and Vincent J. Kalkman and Dan Stowell and A. Leonie Baier},
|
| 102 |
-
year={2026},
|
| 103 |
-
eprint={2608.18191},
|
| 104 |
-
archivePrefix={arXiv},
|
| 105 |
-
primaryClass={cs.LG},
|
| 106 |
-
url={https://arxiv.org/abs/2608.18191},
|
| 107 |
-
}
|
| 108 |
-
```
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
library_name: pytorch
|
|
|
|
| 10 |
- perch
|
| 11 |
---
|
| 12 |
|
| 13 |
+
# ⚠️ Deprecated — see [`perchv2-pytorch`](https://github.com/bghani/perchv2-pytorch) instead
|
| 14 |
|
| 15 |
+
**These weights are no longer the recommended way to use this project.** The main repo now provides a PyTorch backbone built by converting Google's actual released ONNX graph directly (via `onnx2torch`) — confirmed **1.00000000 cosine similarity** against `onnxruntime`'s own output, and fully differentiable. That approach needs the original `perch_v2.onnx` file (see the [main repo README](https://github.com/bghani/perchv2-pytorch#getting-the-weights)), not the weights hosted here.
|
| 16 |
|
| 17 |
+
The weights on this page are kept only as `legacy/` in that repo, for reference — **not recommended for new use.** See why below.
|
| 18 |
|
| 19 |
## What this is
|
| 20 |
|
| 21 |
+
Unofficial, community-converted PyTorch weights for Google's [Perch v2](https://www.kaggle.com/models/google/bird-vocalization-classifier) bioacoustic foundation model, built by hand-reconstructing the architecture in `timm`'s `tf_efficientnet_b3` and copying over weights converted from the original JAX/Flax SavedModel. Not produced or endorsed by Google.
|
| 22 |
+
|
| 23 |
+
- **Architecture:** `timm` `tf_efficientnet_b3`, single-channel input (`in_chans=1`), classification head stripped — returns a pooled **1536-dim** embedding.
|
| 24 |
- **File:** `perch_v2_backbone_timm.pt` — backbone weights only (no classification head).
|
| 25 |
|
| 26 |
+
## Known limitation — why this is deprecated
|
| 27 |
+
|
| 28 |
+
Earlier validation reported ~0.80 whole-network cosine similarity, attributed to numerical error accumulating across blocks. That explanation was incomplete. Further investigation found and fixed two real architectural bugs (a stem padding mismatch and a missing convolution bias, both confirmed directly against the ONNX graph), bringing cosine similarity up to **~0.97**.
|
| 29 |
|
| 30 |
+
That still wasn't the full picture. A later three-way comparison — real ONNX output, the `onnx2torch`-converted backbone, and this `timm` backbone — found this backbone sitting at a **relative L2 error of ~0.23–0.29** against the true model, despite the reassuring-looking ~0.97 cosine similarity. Cosine similarity measures an angle between vectors; it can look close to 1.0 while the actual magnitude of the error stays large. This showed up concretely: a linear probe trained on this backbone's frozen embeddings converged slower and to a lower accuracy than one trained on the ONNX-converted backbone's embeddings, on the same task.
|
| 31 |
|
| 32 |
+
A further architectural bug was identified (blocks 5, 8, and 18 also use asymmetric padding, the same class of issue as the original stem bug) but attempting to fix it caused an unexplained regression, so it was left unfixed. Whether this is the source of the remaining L2 gap is an open question — full details in [`legacy/README.md`](https://github.com/bghani/perchv2-pytorch/blob/main/legacy/README.md).
|
| 33 |
|
| 34 |
+
**If you need accurate, trainable embeddings, use the ONNX-converted backbone in the main repo instead — not these weights.**
|
| 35 |
|
| 36 |
+
## Usage (legacy — see the note above)
|
| 37 |
+
|
| 38 |
+
`legacy` is not part of the installed `perchv2_pytorch` package — clone the [main repo](https://github.com/bghani/perchv2-pytorch) and add its root to `sys.path`:
|
| 39 |
|
| 40 |
```python
|
| 41 |
+
import sys
|
| 42 |
+
sys.path.insert(0, "/path/to/perchv2-pytorch")
|
| 43 |
+
|
| 44 |
import torch
|
| 45 |
from huggingface_hub import hf_hub_download
|
| 46 |
+
from legacy import PerchFrontend, Perch2Classifier, Perch2Embedder
|
| 47 |
|
| 48 |
weights_path = hf_hub_download(repo_id="bghani/perch2-pytorch-weights", filename="perch_v2_backbone_timm.pt")
|
| 49 |
|
| 50 |
+
# 1. Frozen features
|
| 51 |
embedder = Perch2Embedder(weights_path=weights_path)
|
| 52 |
embedder.eval()
|
|
|
|
| 53 |
waveform = torch.zeros(4, 160_000) # 5s clips @ 32kHz, batch of 4
|
| 54 |
with torch.no_grad():
|
| 55 |
embeddings = embedder(waveform) # (4, 1536)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
|
| 57 |
+
# 2. Linear probing
|
| 58 |
mel = PerchFrontend()
|
| 59 |
+
model = Perch2Classifier(num_classes=42, mel=mel, weights_path=weights_path, mode="linear_probe")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
+
# 3. Full fine-tuning
|
| 62 |
+
model = Perch2Classifier(num_classes=42, mel=mel, weights_path=weights_path, mode="finetune")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
```
|
| 64 |
|
|
|
|
|
|
|
| 65 |
## License
|
| 66 |
|
| 67 |
Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://github.com/bghani/perchv2-pytorch/blob/main/NOTICE) in the main repo for the full derivative-work attribution.
|
|
|
|
| 69 |
## Citation
|
| 70 |
|
| 71 |
If you use this in published work, please cite the original Perch v2 paper:
|
|
|
|
| 72 |
```
|
| 73 |
@article{van2025perch,
|
| 74 |
+
title={Perch 2.0: The bittern lesson for bioacoustics},
|
| 75 |
+
author={van Merri{"e}nboer, Bart and Dumoulin, Vincent and Hamer, Jenny and Harrell, Lauren and Burns, Andrea and Denton, Tom},
|
| 76 |
+
journal={arXiv preprint arXiv:2508.04665},
|
| 77 |
+
year={2025}
|
| 78 |
}
|
| 79 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|