karstenskyt commited on
Commit
ff0402d
·
verified ·
1 Parent(s): ac34a29

docs: EU AI Act — Intended Use and Non-Use stanza (SEC1 gap analysis)

Browse files
Files changed (1) hide show
  1. README.md +251 -243
README.md CHANGED
@@ -1,243 +1,251 @@
1
- ---
2
- license: cc-by-nc-4.0
3
- language: en
4
- library_name: pytorch
5
- tags:
6
- - soccer-analytics
7
- - player-embeddings
8
- - deep-sets
9
- - 360-data
10
- - transformer
11
- - adversarial-training
12
- - gradient-reversal
13
- datasets:
14
- - luxury-lakehouse/football2vec-360-training-data
15
- - luxury-lakehouse/football2vec-360-embeddings
16
- metrics:
17
- - mlm_accuracy
18
- pipeline_tag: feature-extraction
19
- ---
20
-
21
- # Football2Vec 360-Enriched — Transformer + Deep Sets Player Embeddings
22
-
23
- 144-dimensional player embedding vectors from a 4-layer transformer encoder augmented with a Deep Sets context encoder (Zaheer et al. 2017) trained on **~2M SPADL actions** with StatsBomb 360 freeze-frame data from **323 professional soccer matches**. Adversarial team debiasing via gradient reversal (Ganin et al. 2016) removes team-identity confounds, producing style representations that generalize across teams.
24
-
25
- This model occupies a **separate embedding space** from [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2) (128-dim, event-only). The 360-enriched vectors are not directly comparable to v2 vectors and should not be mixed in downstream similarity search without re-indexing.
26
-
27
- Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
28
-
29
- ## Model Description
30
-
31
- Football2Vec 360-Enriched extends the v2 transformer architecture with a Deep Sets encoder that processes the spatial positions of all visible opponents and teammates at the moment of each action. The 144-dim output captures both individual action sequences and the spatial context in which those actions occur — richer representations than event-only models for players who frequently appear in 360-annotated matches.
32
-
33
- ### Architecture
34
-
35
- | Component | Detail |
36
- |-----------|--------|
37
- | **Token embedding** | 23 SPADL action types → 128d lookup table |
38
- | **Spatial encoding** | MLP(x) + MLP(y) → 128d each, summed with token embedding |
39
- | **Positional embedding** | Learnable, max 512 tokens |
40
- | **Encoder** | 4-layer TransformerEncoder, 4 attention heads, GELU activation, 4x FFN |
41
- | **Pooling** | Mean pooling over valid (non-padding) tokens → 128d |
42
- | **Deep Sets encoder** | Per-player MLP on freeze-frame (x, y, team) → sum-pool → 16d |
43
- | **Output** | Concatenation [128d transformer \|\| 16d Deep Sets] → 144d |
44
- | **Adversarial head** | Gradient reversal layer (λ=0.2) + team classifier |
45
-
46
- ### Two-Stage Training
47
-
48
- **Stage 1 — Masked Language Modeling:** 15% of action tokens are masked; the model predicts the original action type from surrounding context. The Deep Sets encoder processes freeze-frame coordinates at the masked token position, providing spatial context to the transformer.
49
-
50
- **Stage 2 — Adversarial Debiasing:** A team classifier head is attached via a gradient reversal layer (Ganin et al. 2016). The encoder learns to produce embeddings that *cannot* predict which team a player belongs to, removing team-system confounds while retaining individual style signal.
51
-
52
- ### Dual-Vector Architecture
53
-
54
- This model provides the **behavioral** half of a dual-vector player representation:
55
-
56
- | Vector | Dimensions | Source | Captures |
57
- |--------|-----------|--------|----------|
58
- | **Behavioral** (this model) | 144 | Transformer + Deep Sets on SPADL + 360 freeze-frames | Playing style, spatial context, action sequences |
59
- | **Statistical** | 13 | Z-score normalized per-90 stats | Goals, assists, xG, passes, VAEP, defensive metrics |
60
-
61
- Both vectors are stored in PostgreSQL with [pgvector](https://github.com/pgvector/pgvector) HNSW indexes for sub-10ms similarity queries.
62
-
63
- ## Training Data
64
-
65
- | Source | Matches | Events | License |
66
- |--------|---------|--------|---------|
67
- | [StatsBomb 360 Open Data](https://github.com/statsbomb/open-data) | 323 | ~2M | CC-BY 4.0 |
68
-
69
- The 323-match corpus is the complete StatsBomb 360 open-data release. Coverage includes La Liga, Premier League, Champions League, Euro 2020, and Women's World Cup matches with freeze-frame annotations.
70
-
71
- Training data is published as [`luxury-lakehouse/football2vec-360-training-data`](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-training-data) on HF Hub.
72
-
73
- ### Tokenization
74
-
75
- Events are tokenized using the **23-type SPADL vocabulary**: `pass`, `cross`, `throw_in`, `freekick_crossed`, `freekick_short`, `corner_crossed`, `corner_short`, `take_on`, `foul`, `tackle`, `interception`, `shot`, `shot_penalty`, `shot_freekick`, `keeper_save`, `keeper_claim`, `keeper_punch`, `keeper_pick_up`, `clearance`, `bad_touch`, `non_action`, `dribble`, `goalkick`.
76
-
77
- Continuous spatial coordinates (x, y) normalized to [0, 1] on a 105×68m pitch are injected via learned MLP projections. Freeze-frame coordinates encode visible players as an unordered set with a team indicator bit (0 = teammate, 1 = opponent).
78
-
79
- ### Hyperparameters
80
-
81
- | Parameter | Value |
82
- |-----------|-------|
83
- | Hidden dimension | 128 |
84
- | Encoder layers | 4 |
85
- | Attention heads | 4 |
86
- | FFN multiplier | 4x (512) |
87
- | Dropout | 0.1 |
88
- | Max sequence length | 512 |
89
- | MLM mask probability | 0.15 |
90
- | Spatial MLP intermediate dim | 64 |
91
- | Deep Sets MLP dims | [32, 16] |
92
- | Output dimension | 144 |
93
- | Batch size | 256 |
94
- | Learning rate | 1e-4 |
95
- | Weight decay | 0.01 |
96
- | Warmup fraction | 10% |
97
- | Adversarial λ max | 0.2 |
98
- | Adversarial warmup epochs | 5 |
99
-
100
- Training runs on HF Jobs A10G-small GPU (~90 minutes total for both stages).
101
-
102
- ## How to Use
103
-
104
- ### Quick Start
105
-
106
- ```python
107
- from huggingface_hub import hf_hub_download
108
- from safetensors.torch import load_file
109
- import json
110
-
111
- # Download model weights (Stage 2 — adversarial debiased)
112
- weights_path = hf_hub_download("luxury-lakehouse/football2vec-360", "stage2/model.safetensors")
113
- config_path = hf_hub_download("luxury-lakehouse/football2vec-360", "stage2/config.json")
114
-
115
- with open(config_path) as f:
116
- config = json.load(f)
117
-
118
- state_dict = load_file(weights_path)
119
- print(f"Config: {config['hidden_dim']}-dim transformer + {config['deepsets_dim']}-dim Deep Sets")
120
- print(f"Output dimension: {config['output_dim']}") # 144
121
- print(f"Parameters: {sum(p.numel() for p in state_dict.values()):,}")
122
- ```
123
-
124
- ### Pre-Computed Embeddings (recommended)
125
-
126
- For most use cases, load the pre-computed embeddings directly — no model inference needed:
127
-
128
- ```python
129
- from datasets import load_dataset
130
- import numpy as np
131
-
132
- ds = load_dataset("luxury-lakehouse/football2vec-360-embeddings")
133
- df = ds["train"].to_pandas()
134
-
135
- vectors = np.array(df["behavioral_vector"].tolist())
136
- print(f"{vectors.shape[0]} player-matches, {vectors.shape[1]}-dim embeddings") # (~4K, 144)
137
- ```
138
-
139
- > **Note:** These embeddings cover only players with StatsBomb 360 match appearances (~4K player-match records vs. ~87K for Football2Vec v2). For broader coverage, use [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2).
140
-
141
- ## Intended Use
142
-
143
- - **Context-aware player similarity**: Cosine distance on 144-dim vectors finds players with similar style *and* spatial decision-making
144
- - **Spatial pattern analysis**: The 16-dim Deep Sets component captures how players behave relative to nearby opponents and teammates
145
- - **Scouting in high-press contexts**: Embeddings encode how players handle actions under spatial pressure from surrounding defenders
146
- - **Research**: Reproducible 360-enriched player representations with adversarial debiasing; pairs with Football2Vec v2 for ablation studies
147
- - **Downstream features**: Input to GNN tactical models where spatial context matters
148
-
149
- ## Limitations
150
-
151
- - **360-data only**: Covers 323 StatsBomb 360 matches. Players with appearances only in non-360 matches have no embeddings from this model.
152
- - **Smaller training corpus**: 323 matches vs. ~3,000 for Football2Vec v2. Embeddings for players with few 360 appearances may be noisier.
153
- - **Separate embedding space**: 144-dim vectors are not comparable to Football2Vec v2 128-dim vectors. Cannot mix in the same similarity index without re-embedding all players.
154
- - **Event-based actions + freeze-frames**: Off-ball runs and pressing without a nearby action event are not captured.
155
- - **Team debiasing, not competition debiasing**: The adversarial head targets team ID (stronger confounder in the smaller 360 corpus). Cross-league confounds are attenuated but not fully removed.
156
- - **Open data only**: Derived from publicly available StatsBomb 360 data. Commercial datasets with proprietary 360 annotations may yield different representations.
157
-
158
- ## Freshness
159
-
160
- | Metric | Value |
161
- |--------|-------|
162
- | **Training data freshness SLA** | 168 hours (7 days) |
163
- | **Inference schedule** | Daily 06:00 UTC |
164
- | **Skip guard** | `match_id`-level — only new 360 matches are processed |
165
-
166
- ## Model Files
167
-
168
- ```
169
- stage1/model.safetensors -- Stage 1 MLM checkpoint (safetensors format)
170
- stage2/model.safetensors -- Stage 2 adversarial, final (safetensors format)
171
- stage2/config.json -- Football2Vec360Config as JSON
172
- zscore_params.json -- z-score normalization parameters (13-dim stat vector)
173
- ```
174
-
175
- Model weights use the **safetensors** format — a tensor-only serialization with zero pickle surface and no code execution capability. Pre-computed embeddings are delivered as Parquet (non-executable).
176
-
177
- ## Citation
178
-
179
- ```bibtex
180
- @inproceedings{theiner2022explainable,
181
- title={Explainable Expected Goal Models for Performance Analysis in Football},
182
- author={Theiner, Jonas and M{\"u}ller-Budack, Eric and Ewerth, Ralph},
183
- booktitle={Proceedings of the 4th International Workshop on Multimedia Content Analysis in Sports},
184
- pages={39--47},
185
- year={2022}
186
- }
187
- ```
188
-
189
- ```bibtex
190
- @inproceedings{zaheer2017deep,
191
- title={Deep Sets},
192
- author={Zaheer, Manzil and Kottur, Satwik and Ravanbakhsh, Siamak and Poczos, Barnabas and Salakhutdinov, Ruslan and Smola, Alexander},
193
- booktitle={Advances in Neural Information Processing Systems},
194
- volume={30},
195
- year={2017}
196
- }
197
- ```
198
-
199
- ```bibtex
200
- @article{ganin2016domain,
201
- title={Domain-Adversarial Training of Neural Networks},
202
- author={Ganin, Yaroslav and Ustinova, Evgeniya and Cambau, Hana
203
- and Lempitsky, Victor and Laviolette, Fran{\c{c}}ois},
204
- journal={Journal of Machine Learning Research},
205
- volume={17},
206
- number={1},
207
- pages={1--35},
208
- year={2016}
209
- }
210
- ```
211
-
212
- ```bibtex
213
- @software{nielsen2026football2vec_360,
214
- title={Football2Vec 360-Enriched: Transformer + Deep Sets Player Embeddings},
215
- author={Nielsen, Karsten Skytt},
216
- year={2026},
217
- url={https://github.com/karsten-s-nielsen/luxury-lakehouse}
218
- }
219
- ```
220
-
221
- ## Companion Resources
222
-
223
- | Resource | Description |
224
- |----------|-------------|
225
- | [360 Training Data](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-training-data) | SPADL sequences with 360 freeze-frames used for training |
226
- | [360 Player Embeddings](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-embeddings) | Pre-computed 144-dim vectors per player-match |
227
- | [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2) | 128-dim event-only model (~3,000 matches, broader coverage) |
228
- | [Football2Vec v1](https://huggingface.co/luxury-lakehouse/football2vec-statsbomb-wyscout) | 32-dim Doc2Vec baseline |
229
- | [SPADL/VAEP Action Values](https://huggingface.co/datasets/luxury-lakehouse/spadl-vaep-action-values) | Per-action offensive/defensive VAEP valuations |
230
-
231
- ## Demo
232
-
233
- Try the interactive [Soccer Analytics App](https://huggingface.co/spaces/luxury-lakehouse/soccer-analytics-app) — the Player Similarity page supports similarity search on 360-enriched embeddings for players with 360 match coverage.
234
-
235
- > **Explore interactively:** [HF Space demo](https://huggingface.co/spaces/luxury-lakehouse/soccer-analytics-demo)
236
-
237
- ## More Information
238
-
239
- - **License**: [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) (inherited from StatsBomb open data terms)
240
- - **v2 event-only model**: [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2)
241
- - **v1 baseline model**: [Football2Vec v1 (Doc2Vec)](https://huggingface.co/luxury-lakehouse/football2vec-statsbomb-wyscout)
242
- - **Platform**: [Luxury Lakehouse Soccer Analytics](https://github.com/karsten-s-nielsen/luxury-lakehouse)
243
- - **Workflow card**: `workflow-cards/wf-football2vec-360.yaml`
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-nc-4.0
3
+ language: en
4
+ library_name: pytorch
5
+ tags:
6
+ - soccer-analytics
7
+ - player-embeddings
8
+ - deep-sets
9
+ - 360-data
10
+ - transformer
11
+ - adversarial-training
12
+ - gradient-reversal
13
+ datasets:
14
+ - luxury-lakehouse/football2vec-360-training-data
15
+ - luxury-lakehouse/football2vec-360-embeddings
16
+ metrics:
17
+ - mlm_accuracy
18
+ pipeline_tag: feature-extraction
19
+ ---
20
+
21
+ # Football2Vec 360-Enriched — Transformer + Deep Sets Player Embeddings
22
+
23
+ 144-dimensional player embedding vectors from a 4-layer transformer encoder augmented with a Deep Sets context encoder (Zaheer et al. 2017) trained on **~2M SPADL actions** with StatsBomb 360 freeze-frame data from **323 professional soccer matches**. Adversarial team debiasing via gradient reversal (Ganin et al. 2016) removes team-identity confounds, producing style representations that generalize across teams.
24
+
25
+ This model occupies a **separate embedding space** from [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2) (128-dim, event-only). The 360-enriched vectors are not directly comparable to v2 vectors and should not be mixed in downstream similarity search without re-indexing.
26
+
27
+ Part of the (Right! Luxury!) Lakehouse soccer analytics platform.
28
+
29
+ ## Model Description
30
+
31
+ Football2Vec 360-Enriched extends the v2 transformer architecture with a Deep Sets encoder that processes the spatial positions of all visible opponents and teammates at the moment of each action. The 144-dim output captures both individual action sequences and the spatial context in which those actions occur — richer representations than event-only models for players who frequently appear in 360-annotated matches.
32
+
33
+ ### Architecture
34
+
35
+ | Component | Detail |
36
+ |-----------|--------|
37
+ | **Token embedding** | 23 SPADL action types → 128d lookup table |
38
+ | **Spatial encoding** | MLP(x) + MLP(y) → 128d each, summed with token embedding |
39
+ | **Positional embedding** | Learnable, max 512 tokens |
40
+ | **Encoder** | 4-layer TransformerEncoder, 4 attention heads, GELU activation, 4x FFN |
41
+ | **Pooling** | Mean pooling over valid (non-padding) tokens → 128d |
42
+ | **Deep Sets encoder** | Per-player MLP on freeze-frame (x, y, team) → sum-pool → 16d |
43
+ | **Output** | Concatenation [128d transformer \|\| 16d Deep Sets] → 144d |
44
+ | **Adversarial head** | Gradient reversal layer (λ=0.2) + team classifier |
45
+
46
+ ### Two-Stage Training
47
+
48
+ **Stage 1 — Masked Language Modeling:** 15% of action tokens are masked; the model predicts the original action type from surrounding context. The Deep Sets encoder processes freeze-frame coordinates at the masked token position, providing spatial context to the transformer.
49
+
50
+ **Stage 2 — Adversarial Debiasing:** A team classifier head is attached via a gradient reversal layer (Ganin et al. 2016). The encoder learns to produce embeddings that *cannot* predict which team a player belongs to, removing team-system confounds while retaining individual style signal.
51
+
52
+ ### Dual-Vector Architecture
53
+
54
+ This model provides the **behavioral** half of a dual-vector player representation:
55
+
56
+ | Vector | Dimensions | Source | Captures |
57
+ |--------|-----------|--------|----------|
58
+ | **Behavioral** (this model) | 144 | Transformer + Deep Sets on SPADL + 360 freeze-frames | Playing style, spatial context, action sequences |
59
+ | **Statistical** | 13 | Z-score normalized per-90 stats | Goals, assists, xG, passes, VAEP, defensive metrics |
60
+
61
+ Both vectors are stored in PostgreSQL with [pgvector](https://github.com/pgvector/pgvector) HNSW indexes for sub-10ms similarity queries.
62
+
63
+ ## Training Data
64
+
65
+ | Source | Matches | Events | License |
66
+ |--------|---------|--------|---------|
67
+ | [StatsBomb 360 Open Data](https://github.com/statsbomb/open-data) | 323 | ~2M | CC-BY 4.0 |
68
+
69
+ The 323-match corpus is the complete StatsBomb 360 open-data release. Coverage includes La Liga, Premier League, Champions League, Euro 2020, and Women's World Cup matches with freeze-frame annotations.
70
+
71
+ Training data is published as [`luxury-lakehouse/football2vec-360-training-data`](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-training-data) on HF Hub.
72
+
73
+ ### Tokenization
74
+
75
+ Events are tokenized using the **23-type SPADL vocabulary**: `pass`, `cross`, `throw_in`, `freekick_crossed`, `freekick_short`, `corner_crossed`, `corner_short`, `take_on`, `foul`, `tackle`, `interception`, `shot`, `shot_penalty`, `shot_freekick`, `keeper_save`, `keeper_claim`, `keeper_punch`, `keeper_pick_up`, `clearance`, `bad_touch`, `non_action`, `dribble`, `goalkick`.
76
+
77
+ Continuous spatial coordinates (x, y) normalized to [0, 1] on a 105×68m pitch are injected via learned MLP projections. Freeze-frame coordinates encode visible players as an unordered set with a team indicator bit (0 = teammate, 1 = opponent).
78
+
79
+ ### Hyperparameters
80
+
81
+ | Parameter | Value |
82
+ |-----------|-------|
83
+ | Hidden dimension | 128 |
84
+ | Encoder layers | 4 |
85
+ | Attention heads | 4 |
86
+ | FFN multiplier | 4x (512) |
87
+ | Dropout | 0.1 |
88
+ | Max sequence length | 512 |
89
+ | MLM mask probability | 0.15 |
90
+ | Spatial MLP intermediate dim | 64 |
91
+ | Deep Sets MLP dims | [32, 16] |
92
+ | Output dimension | 144 |
93
+ | Batch size | 256 |
94
+ | Learning rate | 1e-4 |
95
+ | Weight decay | 0.01 |
96
+ | Warmup fraction | 10% |
97
+ | Adversarial λ max | 0.2 |
98
+ | Adversarial warmup epochs | 5 |
99
+
100
+ Training runs on HF Jobs A10G-small GPU (~90 minutes total for both stages).
101
+
102
+ ## How to Use
103
+
104
+ ### Quick Start
105
+
106
+ ```python
107
+ from huggingface_hub import hf_hub_download
108
+ from safetensors.torch import load_file
109
+ import json
110
+
111
+ # Download model weights (Stage 2 — adversarial debiased)
112
+ weights_path = hf_hub_download("luxury-lakehouse/football2vec-360", "stage2/model.safetensors")
113
+ config_path = hf_hub_download("luxury-lakehouse/football2vec-360", "stage2/config.json")
114
+
115
+ with open(config_path) as f:
116
+ config = json.load(f)
117
+
118
+ state_dict = load_file(weights_path)
119
+ print(f"Config: {config['hidden_dim']}-dim transformer + {config['deepsets_dim']}-dim Deep Sets")
120
+ print(f"Output dimension: {config['output_dim']}") # 144
121
+ print(f"Parameters: {sum(p.numel() for p in state_dict.values()):,}")
122
+ ```
123
+
124
+ ### Pre-Computed Embeddings (recommended)
125
+
126
+ For most use cases, load the pre-computed embeddings directly — no model inference needed:
127
+
128
+ ```python
129
+ from datasets import load_dataset
130
+ import numpy as np
131
+
132
+ ds = load_dataset("luxury-lakehouse/football2vec-360-embeddings")
133
+ df = ds["train"].to_pandas()
134
+
135
+ vectors = np.array(df["behavioral_vector"].tolist())
136
+ print(f"{vectors.shape[0]} player-matches, {vectors.shape[1]}-dim embeddings") # (~4K, 144)
137
+ ```
138
+
139
+ > **Note:** These embeddings cover only players with StatsBomb 360 match appearances (~4K player-match records vs. ~87K for Football2Vec v2). For broader coverage, use [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2).
140
+
141
+ ## Intended Use
142
+
143
+ - **Context-aware player similarity**: Cosine distance on 144-dim vectors finds players with similar style *and* spatial decision-making
144
+ - **Spatial pattern analysis**: The 16-dim Deep Sets component captures how players behave relative to nearby opponents and teammates
145
+ - **Scouting in high-press contexts**: Embeddings encode how players handle actions under spatial pressure from surrounding defenders
146
+ - **Research**: Reproducible 360-enriched player representations with adversarial debiasing; pairs with Football2Vec v2 for ablation studies
147
+ - **Downstream features**: Input to GNN tactical models where spatial context matters
148
+
149
+ ## EU AI Act — Intended Use and Non-Use
150
+
151
+ This model is published for **research and reproducibility** purposes on public, open-licensed match data. It is **not intended for, not validated for, and not supplied to** any use that would fall within Annex III §4 (Employment, workers management and access to self-employment) of Regulation (EU) 2024/1689 — including recruitment or selection of natural persons, decisions affecting work-related contractual relationships, promotion, termination, task allocation based on individual traits, or the monitoring and evaluation of performance and behaviour of workers for employment decisions. Player similarity search is a canonical scouting workflow, and any deployer is responsible for treating this model as decision-support at most, never as a decision system.
152
+
153
+ Any deployer who wishes to use this model for such a purpose is responsible for performing their own conformity assessment under Article 43, for drawing up the technical documentation required by Article 11 and Annex IV, for implementing the human oversight measures required by Article 14, for declaring accuracy metrics under Article 15, and for ensuring the data governance obligations of Article 10 are met. Note specifically that the training data contains no protected attributes and therefore cannot support the group-fairness audits required by Article 10(2)(g) without ingesting additional personal data. The adversarial team debiasing (Ganin et al. 2016) described above addresses a *confounding* effect (team-system leakage) and is not a substitute for an Article 10 protected-attribute audit.
154
+
155
+ See the [`AI_GOVERNANCE.md`](https://github.com/karsten-s-nielsen/luxury-lakehouse/blob/main/AI_GOVERNANCE.md) gap analysis in the source repository for the project's full risk classification, re-classification triggers, and governance posture.
156
+
157
+ ## Limitations
158
+
159
+ - **360-data only**: Covers 323 StatsBomb 360 matches. Players with appearances only in non-360 matches have no embeddings from this model.
160
+ - **Smaller training corpus**: 323 matches vs. ~3,000 for Football2Vec v2. Embeddings for players with few 360 appearances may be noisier.
161
+ - **Separate embedding space**: 144-dim vectors are not comparable to Football2Vec v2 128-dim vectors. Cannot mix in the same similarity index without re-embedding all players.
162
+ - **Event-based actions + freeze-frames**: Off-ball runs and pressing without a nearby action event are not captured.
163
+ - **Team debiasing, not competition debiasing**: The adversarial head targets team ID (stronger confounder in the smaller 360 corpus). Cross-league confounds are attenuated but not fully removed.
164
+ - **Open data only**: Derived from publicly available StatsBomb 360 data. Commercial datasets with proprietary 360 annotations may yield different representations.
165
+
166
+ ## Freshness
167
+
168
+ | Metric | Value |
169
+ |--------|-------|
170
+ | **Training data freshness SLA** | 168 hours (7 days) |
171
+ | **Inference schedule** | Daily 06:00 UTC |
172
+ | **Skip guard** | `match_id`-level — only new 360 matches are processed |
173
+
174
+ ## Model Files
175
+
176
+ ```
177
+ stage1/model.safetensors -- Stage 1 MLM checkpoint (safetensors format)
178
+ stage2/model.safetensors -- Stage 2 adversarial, final (safetensors format)
179
+ stage2/config.json -- Football2Vec360Config as JSON
180
+ zscore_params.json -- z-score normalization parameters (13-dim stat vector)
181
+ ```
182
+
183
+ Model weights use the **safetensors** format — a tensor-only serialization with zero pickle surface and no code execution capability. Pre-computed embeddings are delivered as Parquet (non-executable).
184
+
185
+ ## Citation
186
+
187
+ ```bibtex
188
+ @inproceedings{theiner2022explainable,
189
+ title={Explainable Expected Goal Models for Performance Analysis in Football},
190
+ author={Theiner, Jonas and M{\"u}ller-Budack, Eric and Ewerth, Ralph},
191
+ booktitle={Proceedings of the 4th International Workshop on Multimedia Content Analysis in Sports},
192
+ pages={39--47},
193
+ year={2022}
194
+ }
195
+ ```
196
+
197
+ ```bibtex
198
+ @inproceedings{zaheer2017deep,
199
+ title={Deep Sets},
200
+ author={Zaheer, Manzil and Kottur, Satwik and Ravanbakhsh, Siamak and Poczos, Barnabas and Salakhutdinov, Ruslan and Smola, Alexander},
201
+ booktitle={Advances in Neural Information Processing Systems},
202
+ volume={30},
203
+ year={2017}
204
+ }
205
+ ```
206
+
207
+ ```bibtex
208
+ @article{ganin2016domain,
209
+ title={Domain-Adversarial Training of Neural Networks},
210
+ author={Ganin, Yaroslav and Ustinova, Evgeniya and Cambau, Hana
211
+ and Lempitsky, Victor and Laviolette, Fran{\c{c}}ois},
212
+ journal={Journal of Machine Learning Research},
213
+ volume={17},
214
+ number={1},
215
+ pages={1--35},
216
+ year={2016}
217
+ }
218
+ ```
219
+
220
+ ```bibtex
221
+ @software{nielsen2026football2vec_360,
222
+ title={Football2Vec 360-Enriched: Transformer + Deep Sets Player Embeddings},
223
+ author={Nielsen, Karsten Skytt},
224
+ year={2026},
225
+ url={https://github.com/karsten-s-nielsen/luxury-lakehouse}
226
+ }
227
+ ```
228
+
229
+ ## Companion Resources
230
+
231
+ | Resource | Description |
232
+ |----------|-------------|
233
+ | [360 Training Data](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-training-data) | SPADL sequences with 360 freeze-frames used for training |
234
+ | [360 Player Embeddings](https://huggingface.co/datasets/luxury-lakehouse/football2vec-360-embeddings) | Pre-computed 144-dim vectors per player-match |
235
+ | [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2) | 128-dim event-only model (~3,000 matches, broader coverage) |
236
+ | [Football2Vec v1](https://huggingface.co/luxury-lakehouse/football2vec-statsbomb-wyscout) | 32-dim Doc2Vec baseline |
237
+ | [SPADL/VAEP Action Values](https://huggingface.co/datasets/luxury-lakehouse/spadl-vaep-action-values) | Per-action offensive/defensive VAEP valuations |
238
+
239
+ ## Demo
240
+
241
+ Try the interactive [Soccer Analytics App](https://huggingface.co/spaces/luxury-lakehouse/soccer-analytics-app) — the Player Similarity page supports similarity search on 360-enriched embeddings for players with 360 match coverage.
242
+
243
+ > **Explore interactively:** [HF Space demo](https://huggingface.co/spaces/luxury-lakehouse/soccer-analytics-demo)
244
+
245
+ ## More Information
246
+
247
+ - **License**: [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) (inherited from StatsBomb open data terms)
248
+ - **v2 event-only model**: [Football2Vec v2](https://huggingface.co/luxury-lakehouse/football2vec-v2)
249
+ - **v1 baseline model**: [Football2Vec v1 (Doc2Vec)](https://huggingface.co/luxury-lakehouse/football2vec-statsbomb-wyscout)
250
+ - **Platform**: [Luxury Lakehouse Soccer Analytics](https://github.com/karsten-s-nielsen/luxury-lakehouse)
251
+ - **Workflow card**: `workflow-cards/wf-football2vec-360.yaml`