Image Feature Extraction
timm
Safetensors
re-identification
metric-learning
wildlife
dinov2
gorilla
Eval Results (legacy)
Instructions to use gorilla-watch/GorillaWatch-DINOv2-Giant with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use gorilla-watch/GorillaWatch-DINOv2-Giant with timm:
import timm model = timm.create_model("hf_hub:gorilla-watch/GorillaWatch-DINOv2-Giant", pretrained=True) - Notebooks
- Google Colab
- Kaggle
File size: 7,996 Bytes
753f0ec e7425f2 753f0ec | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | ---
license: cc-by-4.0
library_name: timm
pipeline_tag: image-feature-extraction
tags:
- image-feature-extraction
- re-identification
- metric-learning
- wildlife
- dinov2
- gorilla
datasets:
- gorilla-watch/Gorilla-SPAC-Wild
model-index:
- name: GorillaWatch-DINOv2-Giant
results:
- task:
type: image-feature-extraction
name: facial gorilla re-identification
dataset:
type: gorilla-watch/Gorilla-SPAC-Wild
name: Gorilla-SPAC-Wild
config: face_with_body
split: test
metrics:
- name: Micro Accuracy
type: accuracy
value: 0.5554
- name: Macro Accuracy
type: accuracy
value: 0.4629
- name: Tracklet Micro Accuracy
type: accuracy
value: 0.6121
- name: Tracklet Macro Accuracy
type: accuracy
value: 0.4451
- task:
type: image-feature-extraction
name: facial gorilla re-identification
dataset:
type: gorilla-watch/Gorilla-Zoo-Berlin
name: Gorilla-Zoo-Berlin
config: face_with_body
split: test
metrics:
- name: Micro Accuracy
type: accuracy
value: 0.7657
- name: Macro Accuracy
type: accuracy
value: 0.759
- name: Tracklet Micro Accuracy
type: accuracy
value: 0.8218
- name: Tracklet Macro Accuracy
type: accuracy
value: 0.8044
---
# GorillaWatch-DINOv2-Giant
Gorilla re-identification model from **[GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring](https://arxiv.org/abs/2512.07776)** (WACV 2026). Further project details can be found [here](https://gorilla-watch.github.io/).
A `vit_giant_patch14_dinov2.lvd142m` DINOv2 backbone fine-tuned with hard-mining triplet loss on
[Gorilla-SPAC-Wild](https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild), projecting to a
**256-dimensional embedding**. Identification is done by k-NN retrieval against a
gallery of embeddings, not by classification. The model has no fixed identity vocabulary, to enable
generalisation to individuals unseen during training.
| | |
|---|---|
| Backbone | `vit_giant_patch14_dinov2.lvd142m` |
| Input resolution | 518×518 |
| Embedding dimension | 256 |
| Parameters | 1136.9M |
| Training data | Gorilla-SPAC-Wild (`face_with_body`) |
## Preprocessing
> [!IMPORTANT]
> This model does **not** use timm's default DINOv2 transform. It expects a **square resize**
> (which only preserves the aspect ratio when the input images are already squared, which is the case in our datasets) and normalization with **mean = std = 0.5**, not the ImageNet
> statistics reported in the backbone's `default_cfg`. Using timm's default transform produces
> incorrect embeddings.
```python
transforms.Compose([
transforms.Resize((518, 518)),
transforms.ToTensor(),
transforms.Normalize(mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]),
])
```
`modeling.py` in this repository exposes this as `model.get_transform()`.
## Usage
Here we provide a minimal setup to use the model for feature extraction.
Requires `torch`, `timm`, `safetensors`, `huggingface_hub` and `torchvision`.
```python
import sys, torch
from huggingface_hub import snapshot_download
from PIL import Image
# Fetch weights, config and the self-contained modeling.py in one go
local_dir = snapshot_download("gorilla-watch/GorillaWatch-DINOv2-Giant")
sys.path.insert(0, local_dir)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
from modeling import load_model
model = load_model(local_dir, device=device) # already in eval mode
transform = model.get_transform()
image = Image.open("gorilla.png").convert("RGB")
with torch.no_grad():
embedding = model(transform(image).unsqueeze(0).to(model.device)) # (1, 256)
```
`load_model` also accepts the repo id directly (`load_model("gorilla-watch/GorillaWatch-DINOv2-Giant")`) if you would rather not
manage a local directory.
Identity assignment uses **k-NN with k=5 under Euclidean distance** against a gallery of embeddings.
The paper's protocol masks out gallery entries from the same encounter (same camera on the same
date) to avoid trivially easy matches. The full evaluation code can be found in our [GitHub Repo](https://github.com/gorilla-watch/gorillawatch).
## Training
Fine-tuned from the upstream `vit_giant_patch14_dinov2.lvd142m` DINOv2 checkpoint.
| Hyperparameter | Value |
|---|---|
| Loss | Online triplet, hard mining, Euclidean, margin 0.647 |
| Optimizer | AdamW (β=0.9/0.999, ε=1e-7) |
| Learning rate | 1.9e-7, cosine annealing to 1e-7 |
| Batch size | 8 (effective 48 via 6 gradient accumulation steps) |
| Regularization | L2 = 0.0059, L2-SP = 1.3e-5 |
| Epochs | 100 max, best-validation-loss checkpoint retained |
| Precision | AMP (fp16 autocast, fp32 master weights) |
| Seed | 42 |
The code used to train these models can be found in our [Github Repository](https://github.com/gorilla-watch/gorillawatch).
## Results
k-NN retrieval accuracy (k=5, Euclidean distance). Gallery entries from the same encounter (same camera on the same date) are masked out, so every match is made across encounters. Macro accuracy averages over identities and is the harder number: it weights rarely-seen individuals equally with frequently-seen ones.
### In-domain: Gorilla-SPAC-Wild
Test split of [Gorilla-SPAC-Wild](https://huggingface.co/datasets/gorilla-watch/Gorilla-SPAC-Wild), the distribution the model was fine-tuned on.
| Protocol | Micro accuracy | Macro accuracy |
|---|---|---|
| Per image | 0.5554 | 0.4629 |
| Per tracklet (average pooling) | 0.6121 | 0.4451 |
### Out-of-distribution: Gorilla-Zoo-Berlin
[Gorilla-Zoo-Berlin](https://huggingface.co/datasets/gorilla-watch/Gorilla-Zoo-Berlin) is a **zero-shot domain-transfer test**: the model is applied to footage recorded in the Berlin Zoo, with no fine-tuning on it, so enclosure, lighting, camera hardware and the individuals themselves are all unseen. The numbers are still higher, since the amount of individuals is much lower than in the SPAC dataset. This evaluation clearly shows that the model is able to generalize to new, unseen populations.
| Protocol | Micro accuracy | Macro accuracy |
|---|---|---|
| Per image | 0.7657 | 0.7590 |
| Per tracklet (average pooling) | 0.8218 | 0.8044 |
## Provenance
These weights are bit-identical conversions from the `.pth` files created in the training process. They were converted to the `model.safetensors` format for better integration with HuggingFace.
## License
This model is released under the **CC-BY-4.0 License**.
## Citation
```bibtex
@inproceedings{GorillaWatch2026,
title={GorillaWatch: An Automated System for In-the-Wild Gorilla Re-Identification and Population Monitoring},
booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
author={Maximilian Schall and Felix Leonard Knöfel and Noah Elias König and Jan Jonas Kubeler and Maximilian von Klinski and Joan Wilhelm Linnemann and Xiaoshi Liu and Iven Jelle Schlegelmilch and Ole Woyciniuk and Alexandra Schild and Dante Wasmuht and Magdalena Bermejo Espinet and German Illera Basas and Gerard de Melo},
year={2026},
archivePrefix={arXiv},
eprint={2512.07776}
}
```
## Acknowledgements
The project on which this report is based was funded by the Federal Ministry of Research, Technology and Space under the funding code “KI-Servicezentrum Berlin-Brandenburg” 16IS22092. We acknowledge the support of Sabine Plattner African Charities (SPAC) for their funding to this research. We are grateful to Zoo Berlin for their expert assistance and facility access. This collaboration enabled the development of AI tools capable of being deployed in the wild to directly support gorilla conservation. The responsibility for the content of this publication remains with the authors. |