SebasJanampa commited on
Commit
1b3cf39
·
verified ·
1 Parent(s): 2aa5d6e

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +73 -4
README.md CHANGED
@@ -3,9 +3,78 @@ license: apache-2.0
3
  tags:
4
  - model_hub_mixin
5
  - pytorch_model_hub_mixin
 
 
 
6
  ---
7
 
8
- This model has been pushed to the Hub using the [PytorchModelHubMixin](https://huggingface.co/docs/huggingface_hub/package_reference/mixins#huggingface_hub.PyTorchModelHubMixin) integration:
9
- - Code: [More Information Needed]
10
- - Paper: [More Information Needed]
11
- - Docs: [More Information Needed]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  tags:
4
  - model_hub_mixin
5
  - pytorch_model_hub_mixin
6
+ language:
7
+ - en
8
+ pipeline_tag: keypoint-detection
9
  ---
10
 
11
+ # DETRPose-L-COCO
12
+
13
+ DETRPose-L-COCO is a real-time object detection model introduced in the paper [DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints](https://huggingface.co/papers/2506.13027).
14
+
15
+ - **Repository:** [https://github.com/SebastianJanampa/DETRPose](https://github.com/SebastianJanampa/DETRPose)
16
+ - **Paper:** [https://huggingface.co/papers/2506.13027](https://huggingface.co/papers/2506.13027)
17
+
18
+ ## 📝 Model Description
19
+ DETRPose introduces the first real-time end-to-end framework for multi-person pose estimation.
20
+ By leveraging the hybrid encoder from RT-DETR and the lightweight decoder architecture of D-FINE, DETRPose achieves low-latency inference without sacrificing accuracy.
21
+ The model introduces two primary methodological advancements:
22
+
23
+ * **Pose-LQE Layer:** A specialized head designed to improve confidence scores.
24
+ * **Advanced Training Paradigm:** Incorporates Denoising Keypoints and a custom Keypoint Similarity Varifocal loss function, ensuring robust learning and superior localization performance.
25
+
26
+
27
+ | Model | Dataset | AP | #Params | Latency | GFLOPs |
28
+ | :---: | :---: | :---: | :---: | :---: | :---: |
29
+ | **DETRPose-L** | COCO | 72.5 | 32.8 M | 9.50 ms | 107.1 |
30
+
31
+
32
+ ## 🚀 Installation
33
+
34
+ To use this model, you need to install the inference-ready branch of the DETRPose repository.
35
+ You can directly install the inference-ready branch using `pip`:
36
+ ```shell
37
+ pip install git+https://github.com/SebastianJanampa/DETRPose.git@inference_only
38
+ ````
39
+
40
+ ## 💻 Usage
41
+ This branch is designed to be easy to use for inference.
42
+ Here is a quick example of how to load a model and run it on a live webcam feed.
43
+ ```python
44
+ from detrpose import DETR
45
+
46
+ # Initialization
47
+ model = DETR(model='detrpose_hgnetv2_l')
48
+
49
+ # Inference
50
+ model(source=0) # inference on a webcam
51
+ ```
52
+
53
+
54
+ ## 📜 Citation
55
+ If you use `DETRPose` or its methods in your work, please cite the following BibTeX entries:
56
+
57
+ ```latex
58
+ @misc{janampa2025detrpose,
59
+ title={DETRPose: Real-time end-to-end transformer model for multi-person pose estimation},
60
+ author={Sebastian Janampa and Marios Pattichis},
61
+ year={2025},
62
+ eprint={2506.13027},
63
+ archivePrefix={arXiv},
64
+ primaryClass={cs.CV},
65
+ url={https://arxiv.org/abs/2506.13027},
66
+ }
67
+ ```
68
+
69
+ ## 🙏 Acknowledgement
70
+ This work was supported in part by [Lambda.ai](https://lambda.ai).
71
+
72
+ Our work is built upon [DEIM](https://github.com/Intellindust-AI-Lab/DEIM/tree/main), [D-FINE](https://github.com/Peterande/D-FINE), [Detectron2](https://github.com/facebookresearch/detectron2/tree/main), and [GroupPose](https://github.com/Michel-liu/GroupPose/tree/main).
73
+
74
+ ✨ Feel free to reach out if you have any questions! ✨
75
+
76
+ <div align="left">
77
+ <a href="https://lambda.ai" target="_blank">
78
+ <img src="./assets/lambda_logo2.png" width=500 >
79
+ </a>
80
+ </div>