Add README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: onnx
|
| 4 |
+
pipeline_tag: depth-estimation
|
| 5 |
+
tags:
|
| 6 |
+
- depth-anything
|
| 7 |
+
- onnx
|
| 8 |
+
- depth-estimation
|
| 9 |
+
- visionserve
|
| 10 |
+
base_model: depth-anything/Depth-Anything-V2-Small-hf
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Depth Anything V2 Small — ONNX with dynamic input size
|
| 14 |
+
|
| 15 |
+
ONNX export of [`depth-anything/Depth-Anything-V2-Small-hf`](https://huggingface.co/depth-anything/Depth-Anything-V2-Small-hf)
|
| 16 |
+
(Apache-2.0) for [VisionServe](https://hub.docker.com/r/mtbui2010/visionserve):
|
| 17 |
+
|
| 18 |
+
```bash
|
| 19 |
+
visionserve pull depth-anything-v2
|
| 20 |
+
```
|
| 21 |
+
|
| 22 |
+
Only the **Small** variant is published: Depth Anything V2 Base and Large are CC-BY-NC-4.0.
|
| 23 |
+
|
| 24 |
+
## Why another export
|
| 25 |
+
|
| 26 |
+
The input is `[1, 3, H, W]` with **dynamic H and W**, so the image can be fed at the size the
|
| 27 |
+
reference `DPTImageProcessor` uses (keep the aspect ratio, scale by whichever of 518/W, 518/H is
|
| 28 |
+
closer to 1, each side rounded to a multiple of 14, bicubic). For example, an 848×480 photo is
|
| 29 |
+
fed at 910×518 instead of being squashed to 518×518.
|
| 30 |
+
|
| 31 |
+
A plain dynamic-axes export of this model is wrong. The legacy tracer bakes the 518×518 output
|
| 32 |
+
size into the head (`int(patch_h * 14)` in `modeling_depth_anything.py`), so every input returns
|
| 33 |
+
a 518×518 map. This export keeps that size traced, and it is checked against PyTorch at three
|
| 34 |
+
input shapes (518×518, 910×518 and 518×686).
|
| 35 |
+
|
| 36 |
+
## I/O contract
|
| 37 |
+
|
| 38 |
+
| | name | shape | |
|
| 39 |
+
|---|---|---|---|
|
| 40 |
+
| input | `pixel_values` | `[1, 3, H, W]` f32 | H, W multiples of 14; /255, ImageNet mean/std |
|
| 41 |
+
| output | `predicted_depth` | `[1, H, W]` f32 | relative inverse depth (disparity), unnormalised |
|
| 42 |
+
|
| 43 |
+
## Verification (VisionServe converter, 7 non-square photos)
|
| 44 |
+
|
| 45 |
+
| check | result |
|
| 46 |
+
|---|---|
|
| 47 |
+
| ONNX vs PyTorch, 3 input shapes | max \|Δ\|/scale 1.05e-05 |
|
| 48 |
+
| served preprocessing vs `DPTImageProcessor` | mean 0.25 grey levels, same shapes on 7/7 images |
|
| 49 |
+
| served depth map vs `predicted_depth` | Pearson r min 0.9999, mean 1.000 |
|
| 50 |
+
|
| 51 |
+
## License
|
| 52 |
+
|
| 53 |
+
Apache-2.0, inherited from `depth-anything/Depth-Anything-V2-Small-hf`. The weights are not modified.
|