DOI: 10.3390/sci8090265 ISSN: 2413-4155

A Validated, Population-Scale Dataset of 2.7 Million IRONMAN® Triathlon Records with Separated Transition Times (2002–2026)

Aldo Seffrin, Pantelis Theodoros Nikolaidis, Marilia Santos Andrade, Elias Villiger, Thomas Rosemann, Katja Weiss, Daniel Ferreira, Beat Knechtle

Long-distance triathlon has become a major model for studying endurance performance. Research has nonetheless been constrained by distance-specific datasets that separate full-distance from half-distance events and largely omit transition times. Our aim was to assemble, validate, and describe an openly available dataset that removes both constraints. We compiled 2,706,922 athlete-race records from more than 1500 IRONMAN® and IRONMAN® 70.3 events held between 2002 and 2026, the final season partial, drawn from the official IRONMAN® results platform (75.4%) and a third-party aggregator covering event series absent from it (24.6%). Each record contains swimming, first transition (T1), cycling, second transition (T2), running, and overall times in seconds, with finish status, age group, gender, country, and qualification points. The dataset showed close correspondence between sources and high internal consistency. Split times were identical to the second for 98.8% of 4,732,776 athlete–discipline pairs across all 559 race-years present in both sources, summed splits matched reported overall times within one second for 98.3% of records, and separated T1 and T2 times were available for 84.2% and 84.1% of records—to our knowledge, the first population-level provision of this information. Openly deposited with reproducible collection scripts, the dataset supports research on pacing, transition efficiency, sex differences, non-completion, and qualification equity.