Coddieharsh commited on
Commit
2b1dbe4
·
verified ·
1 Parent(s): 55afccb

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +102 -0
README.md ADDED
@@ -0,0 +1,102 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: research-use-only
4
+ license_link: https://github.com/yastrebksv/TennisCourtDetector
5
+ tags:
6
+ - computer-vision
7
+ - keypoint-detection
8
+ - tennis
9
+ - sports-analytics
10
+ ---
11
+
12
+ # Tennis court keypoint detector — geometrically fine-tuned
13
+
14
+ ResNet-50 regression head predicting the **14 standard tennis court keypoints** from a
15
+ single broadcast frame. Used by [Tennis-Vision](https://github.com/HarshTomar1234/Tennis-Vision)
16
+ to fit the court homography that every real-world measurement depends on.
17
+
18
+ ## Why this fine-tune exists
19
+
20
+ The base model is accurate on footage resembling its training distribution but
21
+ degrades on real broadcast clips with different camera framing. Measured on a 9-clip
22
+ YouTube evaluation suite, only about 4 of 9 clips produced a usable court fit — and
23
+ the failures were not obvious, because the model returns a tidy quadrilateral even
24
+ when it lands on the stands.
25
+
26
+ Per-surface accuracy was measured first and ruled out the obvious hypothesis:
27
+
28
+ | surface | median keypoint error | within 25 px |
29
+ |---|---|---|
30
+ | hard (blue) | 3.90 px | 98.5 % |
31
+ | clay | 4.58 px | 98.0 % |
32
+ | grass / green | 4.65 px | 98.2 % |
33
+
34
+ All three within 0.75 px of each other, so surface was never the weakness. Every
35
+ observed real-world failure was the predicted court displaced **vertically**.
36
+ Augmentation is therefore **geometric** — translation (wider vertically), scale, mild
37
+ perspective warp, horizontal flip with keypoint remapping — and deliberately **not**
38
+ photometric. Colour jitter would have trained hard and fixed nothing.
39
+
40
+ ## Results
41
+
42
+ Held-out validation split of the TennisCourtDetector dataset (2,211 images):
43
+
44
+ | metric | base | fine-tuned |
45
+ |---|---|---|
46
+ | median keypoint error | 4.03 px | **2.90 px** |
47
+ | mean keypoint error | 5.71 px | **3.92 px** |
48
+ | keypoints within 10 px | 93.7 % | **97.2 %** |
49
+ | images with all 14 keypoints within 25 px | 96.8 % | **98.3 %** |
50
+
51
+ On the real-clip evaluation suite, clips passing the court-validity gate went from
52
+ **4/9 to 8/9**. Validated on Wimbledon grass footage the model had never seen
53
+ (line-support 0.65 / 0.64 / 0.63). The one remaining failure is a near-ground-level
54
+ camera where the court is extremely foreshortened; the pipeline's validity gate flags
55
+ that case rather than reporting confident wrong numbers.
56
+
57
+ Both numbers are reproducible:
58
+
59
+ ```bash
60
+ python eval/court_keypoint_accuracy.py --model models/keypoints_model_geoaug.pth
61
+ python eval/court_validity_calibration.py --model models/keypoints_model_geoaug.pth
62
+ ```
63
+
64
+ ## Usage
65
+
66
+ ```python
67
+ import torch, cv2
68
+ from torchvision import models
69
+
70
+ model = models.resnet50(weights=None)
71
+ model.fc = torch.nn.Linear(model.fc.in_features, 14 * 2)
72
+ model.load_state_dict(torch.load("keypoints_model_geoaug.pth", map_location="cpu"))
73
+ model.eval()
74
+ # Input: 224x224 RGB, ImageNet normalisation.
75
+ # Output: 28 values (x0, y0, ... x13, y13) in 224x224 space — rescale by
76
+ # original_width / 224 and original_height / 224.
77
+ ```
78
+
79
+ Keypoint order (index pairs forming court lines): `(0,2)` left outer sideline,
80
+ `(1,3)` right outer sideline, `(0,1)` far baseline, `(2,3)` near baseline,
81
+ `(4,5)` left singles sideline, `(6,7)` right singles sideline, `(8,9)` far service
82
+ line, `(10,11)` near service line, `(12,13)` centre service line.
83
+
84
+ ## Limitations
85
+
86
+ - Trained and evaluated on **broadcast and elevated fixed-camera** footage. Ground-level
87
+ cameras fail.
88
+ - Doubles and amateur footage are untested.
89
+ - The validation split shares lineage with the base model's training data, so 2.90 px
90
+ is an in-distribution figure. The real-clip result (8/9) is the out-of-distribution
91
+ evidence.
92
+
93
+ ## Provenance and credit
94
+
95
+ Fine-tuned from the pretrained weights in
96
+ **[yastrebksv/TennisCourtDetector](https://github.com/yastrebksv/TennisCourtDetector)**,
97
+ on that project's dataset (8,841 annotated images across hard, clay and grass). All
98
+ credit for the original architecture, dataset and base weights belongs to that author.
99
+
100
+ The upstream licence is not explicitly stated; the dataset was released for research
101
+ reproduction. This derivative is published on the same terms — **research use only** —
102
+ and should not be assumed to carry any broader grant.