File size: 7,996 Bytes
753f0ec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e7425f2
753f0ec
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
---
license: cc-by-4.0
library_name: timm
pipeline_tag: image-feature-extraction
tags:
  - image-feature-extraction
  - re-identification
  - metric-learning
  - wildlife
  - dinov2
  - gorilla
datasets:
  - gorilla-watch/Gorilla-SPAC-Wild
model-index:
- name: GorillaWatch-DINOv2-Giant
  results:
    - task:
        type: image-feature-extraction
        name: facial gorilla re-identification
      dataset:
        type: gorilla-watch/Gorilla-SPAC-Wild
        name: Gorilla-SPAC-Wild
        config: face_with_body
        split: test
      metrics:
        - name: Micro Accuracy
          type: accuracy
          value: 0.5554
        - name: Macro Accuracy
          type: accuracy
          value: 0.4629
        - name: Tracklet Micro Accuracy
          type: accuracy
          value: 0.6121
        - name: Tracklet Macro Accuracy
          type: accuracy
          value: 0.4451
    - task:
        type: image-feature-extraction
        name: facial gorilla re-identification
      dataset:
        type: gorilla-watch/Gorilla-Zoo-Berlin
        name: Gorilla-Zoo-Berlin
        config: face_with_body
        split: test
      metrics:
        - name: Micro Accuracy
          type: accuracy
          value: 0.7657
        - name: Macro Accuracy
          type: accuracy
          value: 0.759
        - name: Tracklet Micro Accuracy
          type: accuracy
          value: 0.8218
        - name: Tracklet Macro Accuracy
          type: accuracy
          value: 0.8044
---

# GorillaWatch-DINOv2-Giant

Gorilla re-identification model from **[GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring](https://arxiv.org/abs/2512.07776)** (WACV 2026). Further project details can be found [here](https://gorilla-watch.github.io/).

A `vit_giant_patch14_dinov2.lvd142m` DINOv2 backbone fine-tuned with hard-mining triplet loss on
[Gorilla-SPAC-Wild](https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild), projecting to a
**256-dimensional embedding**. Identification is done by k-NN retrieval against a
gallery of embeddings, not by classification. The model has no fixed identity vocabulary, to enable
generalisation to individuals unseen during training.

| | |
|---|---|
| Backbone | `vit_giant_patch14_dinov2.lvd142m` |
| Input resolution | 518×518 |
| Embedding dimension | 256 |
| Parameters | 1136.9M |
| Training data | Gorilla-SPAC-Wild (`face_with_body`) |

## Preprocessing

> [!IMPORTANT]
> This model does **not** use timm's default DINOv2 transform. It expects a **square resize**
> (which only preserves the aspect ratio when the input images are already squared, which is the case in our datasets) and normalization with **mean = std = 0.5**, not the ImageNet
> statistics reported in the backbone's `default_cfg`. Using timm's default transform produces
> incorrect embeddings.

```python
transforms.Compose([
    transforms.Resize((518, 518)),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),
])
```

`modeling.py` in this repository exposes this as `model.get_transform()`.

## Usage

Here we provide a minimal setup to use the model for feature extraction. 

Requires `torch`, `timm`, `safetensors`, `huggingface_hub` and `torchvision`.

```python
import sys, torch
from huggingface_hub import snapshot_download
from PIL import Image

# Fetch weights, config and the self-contained modeling.py in one go
local_dir = snapshot_download("gorilla-watch/GorillaWatch-DINOv2-Giant")
sys.path.insert(0, local_dir)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
from modeling import load_model

model = load_model(local_dir, device=device)   # already in eval mode
transform = model.get_transform()

image = Image.open("gorilla.png").convert("RGB")
with torch.no_grad():
    embedding = model(transform(image).unsqueeze(0).to(model.device))  # (1, 256)
```

`load_model` also accepts the repo id directly (`load_model("gorilla-watch/GorillaWatch-DINOv2-Giant")`) if you would rather not
manage a local directory.

Identity assignment uses **k-NN with k=5 under Euclidean distance** against a gallery of embeddings.
The paper's protocol masks out gallery entries from the same encounter (same camera on the same
date) to avoid trivially easy matches. The full evaluation code can be found in our [GitHub Repo](https://github.com/gorilla-watch/gorillawatch).

## Training

Fine-tuned from the upstream `vit_giant_patch14_dinov2.lvd142m` DINOv2 checkpoint.

| Hyperparameter | Value |
|---|---|
| Loss | Online triplet, hard mining, Euclidean, margin 0.647 |
| Optimizer | AdamW (β=0.9/0.999, ε=1e-7) |
| Learning rate | 1.9e-7, cosine annealing to 1e-7 |
| Batch size | 8 (effective 48 via 6 gradient accumulation steps) |
| Regularization | L2 = 0.0059, L2-SP = 1.3e-5 |
| Epochs | 100 max, best-validation-loss checkpoint retained |
| Precision | AMP (fp16 autocast, fp32 master weights) |
| Seed | 42 |

The code used to train these models can be found in our [Github Repository](https://github.com/gorilla-watch/gorillawatch).

## Results

k-NN retrieval accuracy (k=5, Euclidean distance). Gallery entries from the same encounter (same camera on the same date) are masked out, so every match is made across encounters. Macro accuracy averages over identities and is the harder number: it weights rarely-seen individuals equally with frequently-seen ones.

### In-domain: Gorilla-SPAC-Wild

Test split of [Gorilla-SPAC-Wild](https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild), the distribution the model was fine-tuned on.

| Protocol | Micro accuracy | Macro accuracy |
|---|---|---|
| Per image | 0.5554 | 0.4629 |
| Per tracklet (average pooling) | 0.6121 | 0.4451 |

### Out-of-distribution: Gorilla-Zoo-Berlin

[Gorilla-Zoo-Berlin](https://huggingface.co/datasets/gorilla-watch/Gorilla-Zoo-Berlin) is a **zero-shot domain-transfer test**: the model is applied to footage recorded in the Berlin Zoo, with no fine-tuning on it, so enclosure, lighting, camera hardware and the individuals themselves are all unseen. The numbers are still higher, since the amount of individuals is much lower than in the SPAC dataset. This evaluation clearly shows that the model is able to generalize to new, unseen populations.

| Protocol | Micro accuracy | Macro accuracy |
|---|---|---|
| Per image | 0.7657 | 0.7590 |
| Per tracklet (average pooling) | 0.8218 | 0.8044 |


## Provenance

These weights are bit-identical conversions from the `.pth` files created in the training process. They were converted to the `model.safetensors` format for better integration with HuggingFace.

## License

This model is released under the **CC-BY-4.0 License**.

## Citation

```bibtex
@inproceedings{GorillaWatch2026,
  title={GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring},
  booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
  author={Maximilian Schall and Felix Leonard Knöfel and Noah Elias König and Jan Jonas Kubeler and Maximilian von Klinski and Joan Wilhelm Linnemann and Xiaoshi Liu and Iven Jelle Schlegelmilch and Ole Woyciniuk and Alexandra Schild and Dante Wasmuht and Magdalena Bermejo Espinet and German Illera Basas and Gerard de Melo},
  year={2026},
  archivePrefix={arXiv},
  eprint={2512.07776}
}
```

## Acknowledgements
The project on which this report is based was funded by the Federal Ministry of Research, Technology and Space under the funding code “KI-Servicezentrum Berlin-Brandenburg” 16IS22092. We acknowledge the support of Sabine Plattner African Charities (SPAC) for their funding to this research. We are grateful to Zoo Berlin for their expert assistance and facility access. This collaboration enabled the development of AI tools capable of being deployed in the wild to directly support gorilla conservation. The responsibility for the content of this publication remains with the authors.