bghani commited on
Commit
be786b1
·
verified ·
1 Parent(s): 25a1818

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +28 -57
README.md CHANGED
@@ -1,4 +1,3 @@
1
-
2
  ---
3
  license: apache-2.0
4
  library_name: pytorch
@@ -11,72 +10,58 @@ tags:
11
  - perch
12
  ---
13
 
14
- # Perch v2 — PyTorch backbone weights
15
 
16
- Unofficial, community-converted PyTorch weights for Google's [Perch v2](https://www.kaggle.com/models/google/bird-vocalization-classifier) bioacoustic foundation model. Not produced or endorsed by Google. For use with [`perchv2-pytorch`](https://github.com/bghani/perchv2-pytorch) — a PyTorch port built for full deep fine-tuning, linear probing, and frozen-feature extraction (unlike the official [ONNX/tflite releases](https://huggingface.co/justinchuby/Perch-onnx), which are inference-only and have no autograd graph).
17
 
 
18
 
19
  ## What this is
20
 
21
- - **Architecture:** stock `timm` `tf_efficientnet_b3`, single-channel input (`in_chans=1`), classification head stripped — returns a pooled **1536-dim** embedding.
22
- - **Source:** extracted from Google's original JAX/Flax SavedModel (`infer.graph.variables`) and converted into a PyTorch state dict compatible with `timm`'s `tf_efficientnet_b3`.
 
23
  - **File:** `perch_v2_backbone_timm.pt` — backbone weights only (no classification head).
24
 
25
- ## Known limitation
 
 
26
 
27
- This is an **independent community conversion**, not an official release from Google. Per-block validation against the original (stem, expand blocks, residual blocks, head) showed 0.999+ cosine similarity, but whole-network cosine similarity plateaus around **~0.80**, likely from small numerical error accumulating across the 26 sequential MBConv blocks. Treat these weights as a strong pretrained initialization for fine-tuning, not a bit-exact reproduction of Google's model.
28
 
29
- If you need bit-exact frozen embeddings rather than a trainable backbone, use the [official ONNX/tflite build](https://huggingface.co/justinchuby/Perch-onnx) instead.
30
 
31
- ## Usage
32
 
33
- ### 1. Frozen features
 
 
34
 
35
  ```python
 
 
 
36
  import torch
37
  from huggingface_hub import hf_hub_download
38
- from perchv2_pytorch import Perch2Embedder
39
 
40
  weights_path = hf_hub_download(repo_id="bghani/perch2-pytorch-weights", filename="perch_v2_backbone_timm.pt")
41
 
 
42
  embedder = Perch2Embedder(weights_path=weights_path)
43
  embedder.eval()
44
-
45
  waveform = torch.zeros(4, 160_000) # 5s clips @ 32kHz, batch of 4
46
  with torch.no_grad():
47
  embeddings = embedder(waveform) # (4, 1536)
48
- ```
49
-
50
- ### 2. Linear probing
51
-
52
- ```python
53
- from perchv2_pytorch import PerchFrontend, Perch2Classifier
54
 
 
55
  mel = PerchFrontend()
56
- model = Perch2Classifier(
57
- num_classes=42,
58
- mel=mel,
59
- weights_path=weights_path,
60
- mode="linear_probe", # backbone frozen, only the new head trains
61
- )
62
- ```
63
-
64
- ### 3. Full fine-tuning
65
 
66
- ```python
67
- from perchv2_pytorch import PerchFrontend, Perch2Classifier
68
-
69
- mel = PerchFrontend()
70
- model = Perch2Classifier(
71
- num_classes=42,
72
- mel=mel,
73
- weights_path=weights_path,
74
- mode="finetune", # backbone unfrozen -- this is the whole point of this repo
75
- )
76
  ```
77
 
78
- See the [main repo](https://github.com/bghani/perchv2-pytorch) for full usage — frozen embeddings, linear probing, and full fine-tuning, with runnable examples and a walkthrough notebook.
79
-
80
  ## License
81
 
82
  Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://github.com/bghani/perchv2-pytorch/blob/main/NOTICE) in the main repo for the full derivative-work attribution.
@@ -84,25 +69,11 @@ Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://g
84
  ## Citation
85
 
86
  If you use this in published work, please cite the original Perch v2 paper:
87
-
88
  ```
89
  @article{van2025perch,
90
- title={Perch 2.0: The bittern lesson for bioacoustics},
91
- author={van Merri{\"e}nboer, Bart and Dumoulin, Vincent and Hamer, Jenny and Harrell, Lauren and Burns, Andrea and Denton, Tom},
92
- journal={arXiv preprint arXiv:2508.04665},
93
- year={2025}
94
  }
95
  ```
96
- If you use this specific PyTorch conversion, please also cite [ChiroEcho](https://arxiv.org/abs/2608.18191), where it was first developed and applied (CV4Ecology - ECCV 2026):
97
-
98
- ```
99
- @misc{ghani2026chiroechoextendingautomatedbat,
100
- title={ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy},
101
- author={Burooj Ghani and Welmoed Eversteijn and Milan van Hirtum and Juan Sebastián Cañas and Vincent J. Kalkman and Dan Stowell and A. Leonie Baier},
102
- year={2026},
103
- eprint={2608.18191},
104
- archivePrefix={arXiv},
105
- primaryClass={cs.LG},
106
- url={https://arxiv.org/abs/2608.18191},
107
- }
108
- ```
 
 
1
  ---
2
  license: apache-2.0
3
  library_name: pytorch
 
10
  - perch
11
  ---
12
 
13
+ # ⚠️ Deprecated — see [`perchv2-pytorch`](https://github.com/bghani/perchv2-pytorch) instead
14
 
15
+ **These weights are no longer the recommended way to use this project.** The main repo now provides a PyTorch backbone built by converting Google's actual released ONNX graph directly (via `onnx2torch`) — confirmed **1.00000000 cosine similarity** against `onnxruntime`'s own output, and fully differentiable. That approach needs the original `perch_v2.onnx` file (see the [main repo README](https://github.com/bghani/perchv2-pytorch#getting-the-weights)), not the weights hosted here.
16
 
17
+ The weights on this page are kept only as `legacy/` in that repo, for reference — **not recommended for new use.** See why below.
18
 
19
  ## What this is
20
 
21
+ Unofficial, community-converted PyTorch weights for Google's [Perch v2](https://www.kaggle.com/models/google/bird-vocalization-classifier) bioacoustic foundation model, built by hand-reconstructing the architecture in `timm`'s `tf_efficientnet_b3` and copying over weights converted from the original JAX/Flax SavedModel. Not produced or endorsed by Google.
22
+
23
+ - **Architecture:** `timm` `tf_efficientnet_b3`, single-channel input (`in_chans=1`), classification head stripped — returns a pooled **1536-dim** embedding.
24
  - **File:** `perch_v2_backbone_timm.pt` — backbone weights only (no classification head).
25
 
26
+ ## Known limitation — why this is deprecated
27
+
28
+ Earlier validation reported ~0.80 whole-network cosine similarity, attributed to numerical error accumulating across blocks. That explanation was incomplete. Further investigation found and fixed two real architectural bugs (a stem padding mismatch and a missing convolution bias, both confirmed directly against the ONNX graph), bringing cosine similarity up to **~0.97**.
29
 
30
+ That still wasn't the full picture. A later three-way comparison — real ONNX output, the `onnx2torch`-converted backbone, and this `timm` backbone — found this backbone sitting at a **relative L2 error of ~0.23–0.29** against the true model, despite the reassuring-looking ~0.97 cosine similarity. Cosine similarity measures an angle between vectors; it can look close to 1.0 while the actual magnitude of the error stays large. This showed up concretely: a linear probe trained on this backbone's frozen embeddings converged slower and to a lower accuracy than one trained on the ONNX-converted backbone's embeddings, on the same task.
31
 
32
+ A further architectural bug was identified (blocks 5, 8, and 18 also use asymmetric padding, the same class of issue as the original stem bug) but attempting to fix it caused an unexplained regression, so it was left unfixed. Whether this is the source of the remaining L2 gap is an open question — full details in [`legacy/README.md`](https://github.com/bghani/perchv2-pytorch/blob/main/legacy/README.md).
33
 
34
+ **If you need accurate, trainable embeddings, use the ONNX-converted backbone in the main repo instead — not these weights.**
35
 
36
+ ## Usage (legacy — see the note above)
37
+
38
+ `legacy` is not part of the installed `perchv2_pytorch` package — clone the [main repo](https://github.com/bghani/perchv2-pytorch) and add its root to `sys.path`:
39
 
40
  ```python
41
+ import sys
42
+ sys.path.insert(0, "/path/to/perchv2-pytorch")
43
+
44
  import torch
45
  from huggingface_hub import hf_hub_download
46
+ from legacy import PerchFrontend, Perch2Classifier, Perch2Embedder
47
 
48
  weights_path = hf_hub_download(repo_id="bghani/perch2-pytorch-weights", filename="perch_v2_backbone_timm.pt")
49
 
50
+ # 1. Frozen features
51
  embedder = Perch2Embedder(weights_path=weights_path)
52
  embedder.eval()
 
53
  waveform = torch.zeros(4, 160_000) # 5s clips @ 32kHz, batch of 4
54
  with torch.no_grad():
55
  embeddings = embedder(waveform) # (4, 1536)
 
 
 
 
 
 
56
 
57
+ # 2. Linear probing
58
  mel = PerchFrontend()
59
+ model = Perch2Classifier(num_classes=42, mel=mel, weights_path=weights_path, mode="linear_probe")
 
 
 
 
 
 
 
 
60
 
61
+ # 3. Full fine-tuning
62
+ model = Perch2Classifier(num_classes=42, mel=mel, weights_path=weights_path, mode="finetune")
 
 
 
 
 
 
 
 
63
  ```
64
 
 
 
65
  ## License
66
 
67
  Apache 2.0, inherited from the original Perch v2 release. See [NOTICE](https://github.com/bghani/perchv2-pytorch/blob/main/NOTICE) in the main repo for the full derivative-work attribution.
 
69
  ## Citation
70
 
71
  If you use this in published work, please cite the original Perch v2 paper:
 
72
  ```
73
  @article{van2025perch,
74
+ title={Perch 2.0: The bittern lesson for bioacoustics},
75
+ author={van Merri{"e}nboer, Bart and Dumoulin, Vincent and Hamer, Jenny and Harrell, Lauren and Burns, Andrea and Denton, Tom},
76
+ journal={arXiv preprint arXiv:2508.04665},
77
+ year={2025}
78
  }
79
  ```