vdd-yolo26n-sem / README.md
dronefreak's picture
Add model-index metadata for HF evaluation results widget
2aee32e verified
|
Raw
History Blame Contribute Delete
8.2 kB
---
license: agpl-3.0
pipeline_tag: image-segmentation
library_name: ultralytics
base_model: "Ultralytics/YOLO26"
tags:
- semantic-segmentation
- aerial-imagery
- drone
- vdd
- yolo26
- ultralytics
- pytorch
- computer-vision
datasets:
- RussRobin/VDD
metrics:
- miou
- pixel-accuracy
model-index:
- name: YOLO26n-sem (VDD)
results:
- task:
type: image-segmentation
name: Semantic Segmentation
dataset:
name: VDD
type: vdd
metrics:
- type: mean_iou
value: 73.99
name: mIoU
- type: accuracy
value: 86.32
name: Pixel Accuracy
source:
name: CABiNet eval pipeline
url: https://github.com/dronefreak/CABiNet
---
# YOLO26n-sem Finetuned on VDD
![License](https://img.shields.io/badge/License-AGPL--3.0-1f6feb?style=flat-square) ![Framework](https://img.shields.io/badge/Framework-Ultralytics-6a5acd?style=flat-square) ![Dataset](https://img.shields.io/badge/Dataset-VDD-0aa1a7?style=flat-square) ![mIoU](https://img.shields.io/badge/mIoU-73.99%25-e8a33d?style=flat-square) ![Status](https://img.shields.io/badge/Status-Trained-2ea44f?style=flat-square) ![Maintained](https://img.shields.io/badge/Maintained-yes-17a2b8?style=flat-square)
Fine-tuned YOLO26n semantic segmentation model for aerial drone imagery using the VDD (Varied Drone Dataset) benchmark dataset.
This model is part of the **VDD Semantic Segmentation Model Zoo**, a collection of CABiNet and YOLO26 models trained and evaluated under a common pipeline for aerial semantic segmentation.
<p align="center">
<img src="demo_grid.png" alt="YOLO26n-sem on VDD: Input / Ground Truth / Prediction">
</p>
<sub>Qualitative results on VDD test-split examples — single-scale (imgsz=1024) inference, no TTA.</sub>
---
## Performance
| Metric | Score |
| ------------------- | --------------- |
| mIoU | 73.99 |
| Pixel Accuracy | 86.32 |
| Parameters (M) | 1.63 |
| FLOPs (GFLOPs @ 1024px) | 11.4 |
---
## VDD Model Zoo
| Rank | Model | mIoU (%) | Pixel Acc (%) | Params (M) | FLOPs (GFLOPs) |
| ---- | --------------------- | ------------- | ------------------ | ----------------- | ----------------- |
| 1 | YOLO26x-sem | 78.83 | 89.79 | 40.16 | 430.9 |
| 2 | YOLO26l-sem | 78.57 | 89.68 | 17.87 | 192.4 |
| 3 | CABiNet (MobileNetV3-Large) | 77.76 | 89.57 | 9.17 | 54.8 |
| 4 | YOLO26m-sem | 77.02 | 88.3 | 14.32 | 152.3 |
| 5 | YOLO26s-sem | 76.35 | 88.27 | 6.50 | 44.4 |
| 6 | YOLO26n-sem | 73.99 | 86.32 | 1.63 | 11.4 |
---
## Per-Class IoU (%)
| Class | YOLO26x-sem | YOLO26l-sem | CABiNet (MobileNetV3-Large) | YOLO26m-sem | YOLO26s-sem | YOLO26n-sem |
| --- | --- | --- | --- | --- | --- | --- |
| Other | 64.15 | 65.88 | 66.28 | 61.04 | 60.99 | 57.27 |
| Wall | 69.26 | 70.93 | 65.87 | 70.06 | 67.48 | 64.76 |
| Road | 72.61 | 72.18 | 70.2 | 70.78 | 69.68 | 67.78 |
| Vegetation | 90.14 | 89.68 | 91.06 | 89.59 | 88.63 | 85.01 |
| Vehicle | 70.99 | 68.52 | 73.38 | 68.76 | 66.08 | 62.96 |
| Roof | 89.34 | 87.53 | 86.0 | 84.52 | 85.83 | 84.86 |
| Water | 95.3 | 95.29 | 91.54 | 94.41 | 95.79 | 95.33 |
---
## Evaluation Visualizations
### Per-Class IoU Bar Chart
![IoU Bar Chart](iou_bar_chart.png)
### Confusion Matrix
![Confusion Matrix](confusion_matrix_normalized.png)
### Loss Curves
![Loss Curves](results.png)
---
## Dataset
[VDD (Varied Drone Dataset)](https://github.com/RussRobin/VDD) is a semantic segmentation benchmark for drone imagery spanning varied altitudes, viewpoints, and scenes (urban, rural, and natural), captured at a uniform native resolution of 4000x3000.
### Classes
- Other
- Wall
- Road
- Vegetation
- Vehicle
- Roof
- Water
---
## Usage
### Install Dependencies
```bash
pip install ultralytics huggingface_hub
```
### Load Model from Hugging Face
```python
from huggingface_hub import hf_hub_download
from ultralytics import YOLO
weights = hf_hub_download(
repo_id="dronefreak/vdd-yolo26n-sem",
filename="best.pt"
)
model = YOLO(weights)
```
### Run Inference
```python
results = model.predict(source="image.png", task="semantic", imgsz=1024)
mask = results[0].semantic_mask.cpu().numpy().data # (H, W) class-ID map
```
---
## Training Configuration
| Setting | Value |
| ------------ | ------------------------------------------ |
| Epochs | 150 |
| Image size | 1024 |
| Batch size | 4 |
| Dataset | VDD (converted images/+masks/ format) |
| Framework | Ultralytics YOLO |
| cls_pw (class weighting) | 0.5 |
---
## Official Resources
- **VDD Semantic Segmentation Model Zoo:** https://huggingface.co/collections/dronefreak/vdd-semantic-segmentation-model-zoo
- **CABiNet repository:** https://github.com/dronefreak/CABiNet
- **CABiNet Paper:** https://arxiv.org/abs/2011.00993v2
- **VDD Dataset (Hugging Face):** https://huggingface.co/datasets/RussRobin/VDD
- **VDD Repository:** https://github.com/RussRobin/VDD
- **VDD Paper (arXiv):** https://arxiv.org/abs/2305.13608
- **VDD Published (JVCIR):** https://www.sciencedirect.com/science/article/pii/S1047320325000434
- **Ultralytics YOLO:** https://github.com/ultralytics/ultralytics
- **Ultralytics YOLO26 Paper:** https://arxiv.org/abs/2606.03748
---
## Training Framework
Trained with the [CABiNet repository](https://github.com/dronefreak/CABiNet), which pairs its own real-time segmentation trainer with a parallel Ultralytics YOLO26-sem pipeline — shared dataset tooling, training/eval, and mIoU benchmarking across UAVid, AeroScapes, and VDD. Star the repo if you find these models useful!
---
## Known Limitations
Performance may degrade in:
* Small training set (280 images) — heavier augmentation (mosaic/mixup/copy-paste) offsets this during training, but rare-class generalization may still be limited
* Rare classes (Vehicle, Roof, Water) are underrepresented relative to Vegetation/Road/Wall
* Very high native resolution (4000x3000, uniform) downsampled to the eval imgsz — fine detail on small objects (e.g. vehicles at altitude) can be lost
* Varied altitude/viewpoint scenes (the dataset's defining trait) can shift the domain between training crops and a given inference image
---
## Citation
Please cite the following:
```bibtex
@article{cai2025vdd,
title={Vdd: Varied drone dataset for semantic segmentation},
author={Cai, Wenxiao and Jin, Ke and Hou, Jinyan and Guo, Cong and Wu, Letian and Yang, Wankou},
journal={Journal of Visual Communication and Image Representation},
volume={109},
pages={104429},
year={2025},
publisher={Elsevier}
}
@INPROCEEDINGS{9560977,
author={Kumaar, Saumya and Lyu, Ye and Nex, Francesco and Yang, Michael Ying},
booktitle={2021 IEEE International Conference on Robotics and Automation (ICRA)},
title={CABiNet: Efficient Context Aggregation Network for Low-Latency Semantic Segmentation},
year={2021},
pages={13517-13524},
doi={10.1109/ICRA48506.2021.9560977}
}
@article{Kumaar_Real-time_Semantic_Segmentation_2021,
author = {Kumaar, Saumya and Lyu, Ye and Nex, Francesco and Yang, Michael Ying},
doi = {10.1016/j.isprsjprs.2021.06.006},
journal = {ISPRS Journal of Photogrammetry and Remote Sensing},
pages = {124--134},
title = {{Real-time Semantic Segmentation with Context Aggregation Network}},
url = {https://www.sciencedirect.com/science/article/pii/S0924271621001647},
volume = {178},
year = {2021}
}
@article{jocher2026ultralytics,
title={Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models},
author={Jocher, Glenn and Qiu, Jing and Liu, Mengyu and Lyu, Shuai and Akyon, Fatih Cagatay and Kalfaoglu, Muhammet Esat},
journal={arXiv preprint arXiv:2606.03748},
year={2026}
}
@software{cabinet_uavid_benchmark,
author = {Kumaar, Saumya},
title = {CABiNet: Semantic Segmentation Benchmarking on UAVid (CABiNet vs. YOLO26)},
url = {https://github.com/dronefreak/CABiNet},
year = {2026}
}
```