Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags: [tennis, ball-tracking, tracknet, heatmap, pytorch]
|
| 4 |
+
---
|
| 5 |
+
|
| 6 |
+
# TrackNet tennis ball tracker (mirror)
|
| 7 |
+
|
| 8 |
+
> **Not trained by me.** These are the published weights from
|
| 9 |
+
> [yastrebksv/TrackNet](https://github.com/yastrebksv/TrackNet), mirrored here so
|
| 10 |
+
> a deployment can pull every weight my tennis pipeline needs from one place.
|
| 11 |
+
> All credit for the model and its training belongs to the original author.
|
| 12 |
+
|
| 13 |
+
## What it is
|
| 14 |
+
|
| 15 |
+
`BallTrackerNet`, 10.7M parameters. A VGG-style encoder with a deconv decoder -
|
| 16 |
+
a CNN, not a transformer. It takes **three consecutive frames** stacked into 9
|
| 17 |
+
channels at 360x640 and outputs a per-pixel classification over 256 intensity
|
| 18 |
+
levels; the ball is the argmax.
|
| 19 |
+
|
| 20 |
+
Three frames is the whole point. A tennis ball in broadcast footage is ~10 px
|
| 21 |
+
and motion-blurred into a streak, frequently indistinguishable from a line
|
| 22 |
+
marking in any single frame. Temporal context is what a per-frame detector like
|
| 23 |
+
YOLO structurally cannot use.
|
| 24 |
+
|
| 25 |
+
## Getting a position out of the heatmap
|
| 26 |
+
|
| 27 |
+
Taking the brightest pixel is fragile - one hot pixel on a shoe wins outright.
|
| 28 |
+
Better: threshold the heatmap, run `cv2.HoughCircles`, and take the circle
|
| 29 |
+
nearest the previous frame's ball. That uses blob shape *and* the fact that a
|
| 30 |
+
ball cannot teleport.
|
| 31 |
+
|
| 32 |
+
## Origin
|
| 33 |
+
|
| 34 |
+
TrackNet: Huang, Liao, Chen, Ik, Peng (NCTU Taiwan),
|
| 35 |
+
[arXiv:1907.03698](https://arxiv.org/abs/1907.03698), AVSS 2019. The original
|
| 36 |
+
was Keras; yastrebksv's is the PyTorch reimplementation these weights come from.
|
| 37 |
+
|
| 38 |
+
Trained on the TrackNet tennis dataset - 81 broadcast clips, 10 matches,
|
| 39 |
+
19,835 labelled frames.
|