File size: 6,135 Bytes
76a6855
 
 
 
 
 
 
 
 
 
 
71fd62f
76a6855
 
 
 
 
 
 
 
 
 
71fd62f
76a6855
71fd62f
76a6855
71fd62f
76a6855
71fd62f
76a6855
71fd62f
 
 
 
 
76a6855
71fd62f
76a6855
71fd62f
76a6855
71fd62f
76a6855
71fd62f
 
 
76a6855
71fd62f
76a6855
 
71fd62f
76a6855
 
 
71fd62f
 
76a6855
 
 
 
 
71fd62f
 
76a6855
 
 
71fd62f
76a6855
71fd62f
76a6855
 
71fd62f
76a6855
 
71fd62f
76a6855
71fd62f
 
 
 
 
76a6855
 
71fd62f
 
 
76a6855
 
71fd62f
 
76a6855
71fd62f
 
 
76a6855
 
71fd62f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76a6855
71fd62f
 
 
76a6855
71fd62f
76a6855
71fd62f
76a6855
71fd62f
 
 
 
 
76a6855
71fd62f
76a6855
 
 
71fd62f
76a6855
71fd62f
76a6855
 
 
 
 
 
 
71fd62f
76a6855
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
---
license: apache-2.0
language: en
tags:
  - audio
  - bird
  - bioacoustics
  - wildlife
  - tflite
  - quantized
  - int8
  - float16
  - raspberry-pi
  - edge-ai
  - perch
library_name: tflite
pipeline_tag: audio-classification
datasets:
  - google/bird-vocalization-classifier
base_model: google/bird-vocalization-classifier
---

# Perch V2 — Optimized TFLite Models for Raspberry Pi

Optimized variants of Google's [Perch V2](https://www.kaggle.com/models/google/bird-vocalization-classifier) bird vocalization classifier for **edge deployment on Raspberry Pi and ARM64 devices**.

Three model variants converted directly from the **official Google SavedModel**, each targeting a different performance/quality trade-off.

## Models

| Model | Size | Inference (RPi 5) | Embedding cosine | Top-1 agree | Top-5 agree | Best for |
|-------|------|-------------------|-----------------|-------------|-------------|----------|
| `perch_v2_original.tflite` | 409 MB | 435 ms | baseline | baseline | baseline | Reference / high-RAM devices |
| `perch_v2_fp16.tflite` | 205 MB | 384 ms | 0.9999 | 100% | 99% | **RPi 5 (recommended)** |
| `perch_v2_dynint8.tflite` | 105 MB | 299 ms | 0.9927 | 93% | 90% | **RPi 4 / low-RAM devices** |

> Benchmarked on Raspberry Pi 5 Model B (8GB, Cortex-A76 @ 2.4GHz), 20 real bird recordings from 20 species, 5 runs each, 4 threads.

## Quick Start

### Choose your model

- **RPi 5 (4-8 GB):** Use `perch_v2_fp16.tflite` — near-perfect accuracy, 2x smaller than original
- **RPi 4 (2-4 GB):** Use `perch_v2_dynint8.tflite` — 4x smaller, 31% faster, very good accuracy
- **Desktop / reference:** Use `perch_v2_original.tflite` — exact Google baseline

### Usage

```python
# Works with ai-edge-litert, tflite-runtime, or tensorflow
from ai_edge_litert.interpreter import Interpreter
import numpy as np

model_path = "perch_v2_fp16.tflite"  # or dynint8, or original
interpreter = Interpreter(model_path=model_path, num_threads=4)
interpreter.allocate_tensors()

inp = interpreter.get_input_details()
out = interpreter.get_output_details()

# Input: 5 seconds of audio at 32 kHz
audio = np.zeros((1, 160000), dtype=np.float32)  # replace with real audio
interpreter.set_tensor(inp[0]["index"], audio)
interpreter.invoke()

# Get species logits (14,795 classes)
logits = interpreter.get_tensor(out[3]["index"])[0]
top_species = np.argsort(logits)[-5:][::-1]
```

### Download a single model

```python
from huggingface_hub import hf_hub_download

# Download only the model you need
model_path = hf_hub_download(
    "ernensbjorn/perch-v2-int8-tflite",
    "perch_v2_fp16.tflite"
)
```

## Model Details

### Architecture

- **Backbone:** EfficientNet-B3 (~12M params for embeddings)
- **Classification head:** ~91M params (~101.8M total)
- **Input:** 5.0 seconds @ 32,000 Hz = 160,000 float32 samples
- **Outputs:**
  - Index 0: Spatial embeddings (16 x 4 x 1536)
  - Index 1: Temporal features
  - Index 2: 1536-dim global embedding
  - Index 3: **14,795 species logits** (use this for classification)

### Species Coverage

~10,340 bird species + frogs, insects, mammals (~14,795 total classes).

Use the included `labels.txt` for class names and `bird_indices.json` to filter bird-only species.

### Quantization Methods

| Variant | Method | What's quantized | File size reduction |
|---------|--------|-----------------|-------------------|
| **original** | None (float32 baseline) | Nothing | 1x |
| **fp16** | TFLite float16 quantization | Weights stored as float16, dequantized at runtime | 2x smaller |
| **dynint8** | TFLite dynamic range quantization | Weights quantized to int8, activations remain float32 | 4x smaller |

All variants were converted directly from the [official Google Perch V2 SavedModel](https://www.kaggle.com/models/google/bird-vocalization-classifier/tensorFlow2/perch_v2) using `tf.lite.TFLiteConverter` with appropriate optimization flags. No binary patching or post-hoc manipulation.

## Detailed Benchmarks

### Raspberry Pi 5 (8 GB, Cortex-A76 @ 2.4 GHz, 4 threads)

| Model | Size | p50 latency | p95 latency | Embedding cosine (mean) | Embedding cosine (min) | Top-1 | Top-5 |
|-------|------|------------|------------|------------------------|----------------------|-------|-------|
| original | 409 MB | 435 ms | 534 ms | baseline | baseline | baseline | baseline |
| fp16 | 205 MB | 384 ms | 477 ms | 0.999994 | 0.999991 | 100% | 99% |
| dynint8 | 105 MB | 299 ms | 405 ms | 0.992748 | 0.972732 | 93% | 90% |

- **Embedding cosine:** Cosine similarity of the 1536-dim embedding vector vs the float32 baseline. Values > 0.99 indicate negligible quality loss for downstream tasks.
- **Top-1/Top-5 agreement:** How often the quantized model's top predicted species matches the original's prediction.
- Test data: 20 real field recordings from 20 species (Rougegorge familier, Courlis cendré, Grive mauvis, Sarcelle d'hiver, Râle d'eau, etc.)

### Raspberry Pi 4 Estimates

The RPi 4 (Cortex-A72 @ 1.8 GHz) is roughly 2-3x slower than the RPi 5. Expected latencies:

| Model | Estimated p50 | RAM needed |
|-------|---------------|------------|
| original | ~1000-1300 ms | ~500 MB |
| fp16 | ~900-1150 ms | ~300 MB |
| **dynint8** | **~700-900 ms** | **~150 MB** |

For RPi 4 with 2 GB RAM, `dynint8` is strongly recommended.

## Origin

Converted from the official [Google Perch V2 SavedModel](https://www.kaggle.com/models/google/bird-vocalization-classifier/tensorFlow2/perch_v2) (hosted by Google researcher [cgeorgiaw on HuggingFace](https://huggingface.co/cgeorgiaw/Perch)).

Created as part of the [Birdash](https://github.com/ernens/birdash) project — an open-source bird detection dashboard and engine for Raspberry Pi.

## License

Apache 2.0 (same as the original Perch V2 model by Google)

## Citation

If you use these models, please cite the original Perch V2 work:

```bibtex
@article{ghani2023global,
  title={Global birdsong embeddings enable superior transfer learning for bioacoustic classification},
  author={Ghani, Burooj and Denton, Tom and Kahl, Stefan and Klinck, Holger},
  journal={Scientific Reports},
  year={2023}
}
```