mtbui2010 commited on
Commit
a0b697a
·
verified ·
1 Parent(s): 0964d5c

Add README.md

Browse files
Files changed (1) hide show
  1. README.md +53 -0
README.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: onnx
4
+ pipeline_tag: depth-estimation
5
+ tags:
6
+ - depth-anything
7
+ - onnx
8
+ - depth-estimation
9
+ - visionserve
10
+ base_model: depth-anything/Depth-Anything-V2-Small-hf
11
+ ---
12
+
13
+ # Depth Anything V2 Small — ONNX with dynamic input size
14
+
15
+ ONNX export of [`depth-anything/Depth-Anything-V2-Small-hf`](https://huggingface.co/depth-anything/Depth-Anything-V2-Small-hf)
16
+ (Apache-2.0) for [VisionServe](https://hub.docker.com/r/mtbui2010/visionserve):
17
+
18
+ ```bash
19
+ visionserve pull depth-anything-v2
20
+ ```
21
+
22
+ Only the **Small** variant is published: Depth Anything V2 Base and Large are CC-BY-NC-4.0.
23
+
24
+ ## Why another export
25
+
26
+ The input is `[1, 3, H, W]` with **dynamic H and W**, so the image can be fed at the size the
27
+ reference `DPTImageProcessor` uses (keep the aspect ratio, scale by whichever of 518/W, 518/H is
28
+ closer to 1, each side rounded to a multiple of 14, bicubic). For example, an 848×480 photo is
29
+ fed at 910×518 instead of being squashed to 518×518.
30
+
31
+ A plain dynamic-axes export of this model is wrong. The legacy tracer bakes the 518×518 output
32
+ size into the head (`int(patch_h * 14)` in `modeling_depth_anything.py`), so every input returns
33
+ a 518×518 map. This export keeps that size traced, and it is checked against PyTorch at three
34
+ input shapes (518×518, 910×518 and 518×686).
35
+
36
+ ## I/O contract
37
+
38
+ | | name | shape | |
39
+ |---|---|---|---|
40
+ | input | `pixel_values` | `[1, 3, H, W]` f32 | H, W multiples of 14; /255, ImageNet mean/std |
41
+ | output | `predicted_depth` | `[1, H, W]` f32 | relative inverse depth (disparity), unnormalised |
42
+
43
+ ## Verification (VisionServe converter, 7 non-square photos)
44
+
45
+ | check | result |
46
+ |---|---|
47
+ | ONNX vs PyTorch, 3 input shapes | max \|Δ\|/scale 1.05e-05 |
48
+ | served preprocessing vs `DPTImageProcessor` | mean 0.25 grey levels, same shapes on 7/7 images |
49
+ | served depth map vs `predicted_depth` | Pearson r min 0.9999, mean 1.000 |
50
+
51
+ ## License
52
+
53
+ Apache-2.0, inherited from `depth-anything/Depth-Anything-V2-Small-hf`. The weights are not modified.