Direct Comparison Between Loss‐to‐Follow‐Up and Statistical Fragility Is Methodologically Inappropriate, and Fragility Reflects the P Value, Not Trial Robustness: A Simulation Analysis of 300,000 Randomi
Prushoth Vivekanantha, Helena Son, Jeffrey Kay, Satyavenkata Kotipalli, Marc Daniel Bouchard, Kim Madden, Nicole Simunovic, Olufemi R. AyeniPurpose
To assess the relation between the fragility index (FI) and reverse fragility index (RFI) with the minimum number of patients needed to reverse statistical significance (e.g. henceforth termed the lost to follow‐up index (LTFI) and reverse LTFI (R‐LTFI), respectively) and apply machine learning to identify which trial parameters are most important in FI, RFI, and continuous fragility index composition given the nonlinearity of these metrics.
Methods
A total of 300,000 randomized controlled trials (100,000 for each metric) were simulated using common trial parameter value ranges. For FI and RFI, LTFI and R‐LTFI values, respectively, were calculated as the minimum number of patients lost to follow‐up to reverse significance in either direction. Machine learning models were trained to assess the relative importance of P value, sample size, and event numbers in FI and RFI composition and P value, group means, group standard deviations, and group sizes for continuous fragility index composition.
Results
Among their respective cohorts of 100,000, the LTFI and R‐LTFI were greater than FI and RFI in 84.9% and 95.5% of simulated randomized controlled trials, respectively. Random Forest and XGBoost machine learning models had near perfect accuracy in modeling fragility metrics (R 2 ≥ 0.98), justifying its use over standard linear regression models in identifying trial parameters that are most important in determining fragility values. Feature importance analysis showed that the P value accounted for 79.8%, 71.9%, and 64.5% of variability in FI, RFI, and continuous fragility index values.
Conclusions
Fragility metrics are primarily mathematical reflections of standard trial parameters and are heavily driven by the P value. Direct comparisons between fragility metrics and lost to follow‐up are statistically inappropriate because LTFI and R‐LTFI routinely exceed FI and RFI, respectively, and should be avoided in future fragility‐based studies.
Clinical Relevance
Understanding that statistical fragility metrics are mathematical transformations of trial parameters, and not independent measures of robustness, can prevent trial misinterpretation. Additionally, recognizing that the number of patients required to be lost to follow‐up to reverse trial significance often exceeds fragility metrics should discourage inappropriate comparisons in future orthopaedic research.