Depth Anything V2 Small โ ONNX with dynamic input size
ONNX export of depth-anything/Depth-Anything-V2-Small-hf
(Apache-2.0) for VisionServe:
visionserve pull depth-anything-v2
Only the Small variant is published: Depth Anything V2 Base and Large are CC-BY-NC-4.0.
Why another export
The input is [1, 3, H, W] with dynamic H and W, so the image can be fed at the size the
reference DPTImageProcessor uses (keep the aspect ratio, scale by whichever of 518/W, 518/H is
closer to 1, each side rounded to a multiple of 14, bicubic). For example, an 848ร480 photo is
fed at 910ร518 instead of being squashed to 518ร518.
A plain dynamic-axes export of this model is wrong. The legacy tracer bakes the 518ร518 output
size into the head (int(patch_h * 14) in modeling_depth_anything.py), so every input returns
a 518ร518 map. This export keeps that size traced, and it is checked against PyTorch at three
input shapes (518ร518, 910ร518 and 518ร686).
I/O contract
| name | shape | ||
|---|---|---|---|
| input | pixel_values |
[1, 3, H, W] f32 |
H, W multiples of 14; /255, ImageNet mean/std |
| output | predicted_depth |
[1, H, W] f32 |
relative inverse depth (disparity), unnormalised |
Verification (VisionServe converter, 7 non-square photos)
| check | result |
|---|---|
| ONNX vs PyTorch, 3 input shapes | max |ฮ|/scale 1.05e-05 |
served preprocessing vs DPTImageProcessor |
mean 0.25 grey levels, same shapes on 7/7 images |
served depth map vs predicted_depth |
Pearson r min 0.9999, mean 1.000 |
License
Apache-2.0, inherited from depth-anything/Depth-Anything-V2-Small-hf. The weights are not modified.
Model tree for mtbui2010/depth-anything-v2-small-ONNX
Base model
depth-anything/Depth-Anything-V2-Small-hf