Push model using huggingface_hub.
Browse files- README.md +4 -73
- config.json +29 -29
- model.safetensors +2 -2
README.md
CHANGED
|
@@ -3,78 +3,9 @@ license: apache-2.0
|
|
| 3 |
tags:
|
| 4 |
- model_hub_mixin
|
| 5 |
- pytorch_model_hub_mixin
|
| 6 |
-
language:
|
| 7 |
-
- en
|
| 8 |
-
pipeline_tag: keypoint-detection
|
| 9 |
---
|
| 10 |
|
| 11 |
-
#
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
- **Repository:** [https://github.com/SebastianJanampa/DETRPose](https://github.com/SebastianJanampa/DETRPose)
|
| 16 |
-
- **Paper:** [https://huggingface.co/papers/2506.13027](https://huggingface.co/papers/2506.13027)
|
| 17 |
-
|
| 18 |
-
## 📝 Model Description
|
| 19 |
-
DETRPose introduces the first real-time end-to-end framework for multi-person pose estimation.
|
| 20 |
-
By leveraging the hybrid encoder from RT-DETR and the lightweight decoder architecture of D-FINE, DETRPose achieves low-latency inference without sacrificing accuracy.
|
| 21 |
-
The model introduces two primary methodological advancements:
|
| 22 |
-
|
| 23 |
-
* **Pose-LQE Layer:** A specialized head designed to improve confidence scores.
|
| 24 |
-
* **Advanced Training Paradigm:** Incorporates Denoising Keypoints and a custom Keypoint Similarity Varifocal loss function, ensuring robust learning and superior localization performance.
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
| Model | Dataset | AP | #Params | Latency | GFLOPs |
|
| 28 |
-
| :---: | :---: | :---: | :---: | :---: | :---: |
|
| 29 |
-
| **DETRPose-L** | COCO | 72.5 | 32.8 M | 9.50 ms | 107.1 |
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
## 🚀 Installation
|
| 33 |
-
|
| 34 |
-
To use this model, you need to install the inference-ready branch of the DETRPose repository.
|
| 35 |
-
You can directly install the inference-ready branch using `pip`:
|
| 36 |
-
```shell
|
| 37 |
-
pip install git+https://github.com/SebastianJanampa/DETRPose.git@inference_only
|
| 38 |
-
````
|
| 39 |
-
|
| 40 |
-
## 💻 Usage
|
| 41 |
-
This branch is designed to be easy to use for inference.
|
| 42 |
-
Here is a quick example of how to load a model and run it on a live webcam feed.
|
| 43 |
-
```python
|
| 44 |
-
from detrpose import DETR
|
| 45 |
-
|
| 46 |
-
# Initialization
|
| 47 |
-
model = DETR(model='detrpose_hgnetv2_l')
|
| 48 |
-
|
| 49 |
-
# Inference
|
| 50 |
-
model(source=0) # inference on a webcam
|
| 51 |
-
```
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
## 📜 Citation
|
| 55 |
-
If you use `DETRPose` or its methods in your work, please cite the following BibTeX entries:
|
| 56 |
-
|
| 57 |
-
```latex
|
| 58 |
-
@misc{janampa2025detrpose,
|
| 59 |
-
title={DETRPose: Real-time end-to-end transformer model for multi-person pose estimation},
|
| 60 |
-
author={Sebastian Janampa and Marios Pattichis},
|
| 61 |
-
year={2025},
|
| 62 |
-
eprint={2506.13027},
|
| 63 |
-
archivePrefix={arXiv},
|
| 64 |
-
primaryClass={cs.CV},
|
| 65 |
-
url={https://arxiv.org/abs/2506.13027},
|
| 66 |
-
}
|
| 67 |
-
```
|
| 68 |
-
|
| 69 |
-
## 🙏 Acknowledgement
|
| 70 |
-
This work was supported in part by [Lambda.ai](https://lambda.ai).
|
| 71 |
-
|
| 72 |
-
Our work is built upon [DEIM](https://github.com/Intellindust-AI-Lab/DEIM/tree/main), [D-FINE](https://github.com/Peterande/D-FINE), [Detectron2](https://github.com/facebookresearch/detectron2/tree/main), and [GroupPose](https://github.com/Michel-liu/GroupPose/tree/main).
|
| 73 |
-
|
| 74 |
-
✨ Feel free to reach out if you have any questions! ✨
|
| 75 |
-
|
| 76 |
-
<div align="left">
|
| 77 |
-
<a href="https://lambda.ai" target="_blank">
|
| 78 |
-
<img src="./assets/lambda_logo2.png" width=500 >
|
| 79 |
-
</a>
|
| 80 |
-
</div>
|
|
|
|
| 3 |
tags:
|
| 4 |
- model_hub_mixin
|
| 5 |
- pytorch_model_hub_mixin
|
|
|
|
|
|
|
|
|
|
| 6 |
---
|
| 7 |
|
| 8 |
+
This model has been pushed to the Hub using the [PytorchModelHubMixin](https://huggingface.co/docs/huggingface_hub/package_reference/mixins#huggingface_hub.PyTorchModelHubMixin) integration:
|
| 9 |
+
- Code: [More Information Needed]
|
| 10 |
+
- Paper: [More Information Needed]
|
| 11 |
+
- Docs: [More Information Needed]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
config.json
CHANGED
|
@@ -15,7 +15,35 @@
|
|
| 15 |
],
|
| 16 |
"use_lab": false
|
| 17 |
},
|
| 18 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
"_target_": "detrpose.nn.DETRPoseTransformer",
|
| 20 |
"activation": "relu",
|
| 21 |
"cls_no_bias": false,
|
|
@@ -48,34 +76,6 @@
|
|
| 48 |
"two_stage_bbox_embed_share": false,
|
| 49 |
"two_stage_class_embed_share": false,
|
| 50 |
"two_stage_type": "standard"
|
| 51 |
-
},
|
| 52 |
-
"encoder": {
|
| 53 |
-
"_target_": "detrpose.nn.DETRPoseHybridEncoder",
|
| 54 |
-
"act": "silu",
|
| 55 |
-
"depth_mult": 1.0,
|
| 56 |
-
"dim_feedforward": 1024,
|
| 57 |
-
"dropout": 0.0,
|
| 58 |
-
"enc_act": "gelu",
|
| 59 |
-
"eval_spatial_size": [
|
| 60 |
-
640,
|
| 61 |
-
640
|
| 62 |
-
],
|
| 63 |
-
"expansion": 1.0,
|
| 64 |
-
"feat_strides": [
|
| 65 |
-
8,
|
| 66 |
-
16,
|
| 67 |
-
32
|
| 68 |
-
],
|
| 69 |
-
"hidden_dim": 256,
|
| 70 |
-
"in_channels": [
|
| 71 |
-
512,
|
| 72 |
-
1024,
|
| 73 |
-
2048
|
| 74 |
-
],
|
| 75 |
-
"num_heads": 8,
|
| 76 |
-
"num_levels": 3,
|
| 77 |
-
"temperatureH": 20,
|
| 78 |
-
"temperatureW": 20
|
| 79 |
}
|
| 80 |
},
|
| 81 |
"postprocessor": {
|
|
|
|
| 15 |
],
|
| 16 |
"use_lab": false
|
| 17 |
},
|
| 18 |
+
"encoder": {
|
| 19 |
+
"_target_": "detrpose.nn.DETRPoseHybridEncoder",
|
| 20 |
+
"act": "silu",
|
| 21 |
+
"depth_mult": 1.0,
|
| 22 |
+
"dim_feedforward": 1024,
|
| 23 |
+
"dropout": 0.0,
|
| 24 |
+
"enc_act": "gelu",
|
| 25 |
+
"eval_spatial_size": [
|
| 26 |
+
640,
|
| 27 |
+
640
|
| 28 |
+
],
|
| 29 |
+
"expansion": 1.0,
|
| 30 |
+
"feat_strides": [
|
| 31 |
+
8,
|
| 32 |
+
16,
|
| 33 |
+
32
|
| 34 |
+
],
|
| 35 |
+
"hidden_dim": 256,
|
| 36 |
+
"in_channels": [
|
| 37 |
+
512,
|
| 38 |
+
1024,
|
| 39 |
+
2048
|
| 40 |
+
],
|
| 41 |
+
"num_heads": 8,
|
| 42 |
+
"num_levels": 3,
|
| 43 |
+
"temperatureH": 20,
|
| 44 |
+
"temperatureW": 20
|
| 45 |
+
},
|
| 46 |
+
"transformer": {
|
| 47 |
"_target_": "detrpose.nn.DETRPoseTransformer",
|
| 48 |
"activation": "relu",
|
| 49 |
"cls_no_bias": false,
|
|
|
|
| 76 |
"two_stage_bbox_embed_share": false,
|
| 77 |
"two_stage_class_embed_share": false,
|
| 78 |
"two_stage_type": "standard"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
}
|
| 80 |
},
|
| 81 |
"postprocessor": {
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b5c2acd3dde633114ae6fbc9b2e7ef16e01735d0d34928c55e87cefc0566432c
|
| 3 |
+
size 133800348
|