DOI: 10.1177/22150218261474554 ISSN: 2215-020X

Synthetic data generation in sport: A framework for design, privacy, utility, fidelity, and deployment

Paul Yomer Ruiz Pinto, Paul Pao-Yen Wu, John Warmenhoven, Kerrie Mengersen, Divya Mehta

The increasing use of synthetic data in sports science reflects persistent limitations in real-world sport datasets, including small sample sizes, class imbalance, privacy restrictions, and the absence of ground truth. While synthetic data generation has been applied across a wide range of sporting contexts, its use remains heterogeneous, with limited consistency in design, evaluation, and reporting practices. This study synthesises evidence from twelve sport studies to address a central gap in the literature: the lack of an established, domain-specific framework to guide the systematic generation, evaluation, and deployment of synthetic data in sport. Rather than proposing a new generative model, this work introduces a decision-oriented framework that structures synthetic data projects across six interrelated dimensions: objective of use, data structure, generation strategy, domain constraints, utility and fidelity evaluation, and deployment risk. Analysis of the reviewed studies shows that synthetic data are most effective when generation strategies are aligned with sport-specific domain knowledge and clearly defined operational goals. The framework highlights recurring challenges related to domain shift, constrained realism, and class imbalance, and demonstrates how these issues can be addressed through explicit design and evaluation choices.

More from our Archive