Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: research-use-only
|
| 4 |
+
license_link: https://github.com/yastrebksv/TennisCourtDetector
|
| 5 |
+
tags:
|
| 6 |
+
- computer-vision
|
| 7 |
+
- keypoint-detection
|
| 8 |
+
- tennis
|
| 9 |
+
- sports-analytics
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Tennis court keypoint detector — geometrically fine-tuned
|
| 13 |
+
|
| 14 |
+
ResNet-50 regression head predicting the **14 standard tennis court keypoints** from a
|
| 15 |
+
single broadcast frame. Used by [Tennis-Vision](https://github.com/HarshTomar1234/Tennis-Vision)
|
| 16 |
+
to fit the court homography that every real-world measurement depends on.
|
| 17 |
+
|
| 18 |
+
## Why this fine-tune exists
|
| 19 |
+
|
| 20 |
+
The base model is accurate on footage resembling its training distribution but
|
| 21 |
+
degrades on real broadcast clips with different camera framing. Measured on a 9-clip
|
| 22 |
+
YouTube evaluation suite, only about 4 of 9 clips produced a usable court fit — and
|
| 23 |
+
the failures were not obvious, because the model returns a tidy quadrilateral even
|
| 24 |
+
when it lands on the stands.
|
| 25 |
+
|
| 26 |
+
Per-surface accuracy was measured first and ruled out the obvious hypothesis:
|
| 27 |
+
|
| 28 |
+
| surface | median keypoint error | within 25 px |
|
| 29 |
+
|---|---|---|
|
| 30 |
+
| hard (blue) | 3.90 px | 98.5 % |
|
| 31 |
+
| clay | 4.58 px | 98.0 % |
|
| 32 |
+
| grass / green | 4.65 px | 98.2 % |
|
| 33 |
+
|
| 34 |
+
All three within 0.75 px of each other, so surface was never the weakness. Every
|
| 35 |
+
observed real-world failure was the predicted court displaced **vertically**.
|
| 36 |
+
Augmentation is therefore **geometric** — translation (wider vertically), scale, mild
|
| 37 |
+
perspective warp, horizontal flip with keypoint remapping — and deliberately **not**
|
| 38 |
+
photometric. Colour jitter would have trained hard and fixed nothing.
|
| 39 |
+
|
| 40 |
+
## Results
|
| 41 |
+
|
| 42 |
+
Held-out validation split of the TennisCourtDetector dataset (2,211 images):
|
| 43 |
+
|
| 44 |
+
| metric | base | fine-tuned |
|
| 45 |
+
|---|---|---|
|
| 46 |
+
| median keypoint error | 4.03 px | **2.90 px** |
|
| 47 |
+
| mean keypoint error | 5.71 px | **3.92 px** |
|
| 48 |
+
| keypoints within 10 px | 93.7 % | **97.2 %** |
|
| 49 |
+
| images with all 14 keypoints within 25 px | 96.8 % | **98.3 %** |
|
| 50 |
+
|
| 51 |
+
On the real-clip evaluation suite, clips passing the court-validity gate went from
|
| 52 |
+
**4/9 to 8/9**. Validated on Wimbledon grass footage the model had never seen
|
| 53 |
+
(line-support 0.65 / 0.64 / 0.63). The one remaining failure is a near-ground-level
|
| 54 |
+
camera where the court is extremely foreshortened; the pipeline's validity gate flags
|
| 55 |
+
that case rather than reporting confident wrong numbers.
|
| 56 |
+
|
| 57 |
+
Both numbers are reproducible:
|
| 58 |
+
|
| 59 |
+
```bash
|
| 60 |
+
python eval/court_keypoint_accuracy.py --model models/keypoints_model_geoaug.pth
|
| 61 |
+
python eval/court_validity_calibration.py --model models/keypoints_model_geoaug.pth
|
| 62 |
+
```
|
| 63 |
+
|
| 64 |
+
## Usage
|
| 65 |
+
|
| 66 |
+
```python
|
| 67 |
+
import torch, cv2
|
| 68 |
+
from torchvision import models
|
| 69 |
+
|
| 70 |
+
model = models.resnet50(weights=None)
|
| 71 |
+
model.fc = torch.nn.Linear(model.fc.in_features, 14 * 2)
|
| 72 |
+
model.load_state_dict(torch.load("keypoints_model_geoaug.pth", map_location="cpu"))
|
| 73 |
+
model.eval()
|
| 74 |
+
# Input: 224x224 RGB, ImageNet normalisation.
|
| 75 |
+
# Output: 28 values (x0, y0, ... x13, y13) in 224x224 space — rescale by
|
| 76 |
+
# original_width / 224 and original_height / 224.
|
| 77 |
+
```
|
| 78 |
+
|
| 79 |
+
Keypoint order (index pairs forming court lines): `(0,2)` left outer sideline,
|
| 80 |
+
`(1,3)` right outer sideline, `(0,1)` far baseline, `(2,3)` near baseline,
|
| 81 |
+
`(4,5)` left singles sideline, `(6,7)` right singles sideline, `(8,9)` far service
|
| 82 |
+
line, `(10,11)` near service line, `(12,13)` centre service line.
|
| 83 |
+
|
| 84 |
+
## Limitations
|
| 85 |
+
|
| 86 |
+
- Trained and evaluated on **broadcast and elevated fixed-camera** footage. Ground-level
|
| 87 |
+
cameras fail.
|
| 88 |
+
- Doubles and amateur footage are untested.
|
| 89 |
+
- The validation split shares lineage with the base model's training data, so 2.90 px
|
| 90 |
+
is an in-distribution figure. The real-clip result (8/9) is the out-of-distribution
|
| 91 |
+
evidence.
|
| 92 |
+
|
| 93 |
+
## Provenance and credit
|
| 94 |
+
|
| 95 |
+
Fine-tuned from the pretrained weights in
|
| 96 |
+
**[yastrebksv/TennisCourtDetector](https://github.com/yastrebksv/TennisCourtDetector)**,
|
| 97 |
+
on that project's dataset (8,841 annotated images across hard, clay and grass). All
|
| 98 |
+
credit for the original architecture, dataset and base weights belongs to that author.
|
| 99 |
+
|
| 100 |
+
The upstream licence is not explicitly stated; the dataset was released for research
|
| 101 |
+
reproduction. This derivative is published on the same terms — **research use only** —
|
| 102 |
+
and should not be assumed to carry any broader grant.
|