Discrete time survival analysis machine learning models for IVF prognosis across multiple cycles
TingTing Shao, Achilleas Ghinis, Eline Dancet, Johanna Devroe, Felipe Kenji Nakano, Celine VensAbstract
STUDY QUESTION
Can discrete time survival analysis machine learning (ML) methods provide reliable and accurate predictions of cumulative live birth rate (CLBR) for patients both at diagnosis (pre-treatment) and after the first cycle of treatment (post-treatment)?
SUMMARY ANSWER
We developed discrete time survival analysis ML models that provide personalized CLBR estimates of up to six complete IVF cycles with good discriminative performance, especially at the post-treatment phase.
WHAT IS KNOWN ALREADY
The first relevant prognostic models to estimate IVF cumulative success rate across multiple cycles have moderate performance, with C-statistics ranging from 0.7 to 0.73 for both pre-treatment and post-treatment models.
STUDY DESIGN, SIZE, DURATION
This retrospective, single-centre study included a cohort of 1,429 couples who started their IVF treatment at a tertiary center between January 2014 and April 2021 as the development dataset. An additional 613 couples who started treatment at the same center between April 2021 and December 2023 were used as the temporal validation dataset.
PARTICIPANTS/MATERIALS, SETTING, METHODS
Predictive performance was assessed using the following metrics: prediction error (Brier score, scaled Brier score), time-dependent AUC (cumulative/dynamic AUC, mean AUC and time-dependent PR-AUC), and calibration (mean calibration, calibration plot, and calibration slope). Both pre-treatment and post-treatment models predicted the cumulative live birth success rates of up to six complete IVF cycles. Various discrete time survival analysis methods were trained, evaluated, and compared with the baseline models. Discrete time survival analysis models include the ones embedded with a generalized linear model (Logit) and tree-based ensembles (Bayesian Additive Regression Trees (BART), random forests). The baseline models are the refitted/recalibrated pre-treatment and post-treatment McLernon models, considering factors available in our data.
MAIN RESULTS AND THE ROLE OF CHANCE
The discrete time survival analysis models, especially BART, performed better on our data than the baseline models especially at the post-treatment phase. Validation results showed moderate discriminative ability for the pre-treatment model, with a mean AUC 0.69 (95% CI: 0.65, 0.73), and a scaled integrated Brier score (scaled IBS) of 0.13 (13% performance improvement over Kaplan-Meier estimates). For the post-treatment model, discriminative power was strong, with a mean AUC of 0.82 (95% CI: 0.79, 0.85), and a scaled IBS of 0.19. In the BART model, the most important feature was female age, and embryo utilization rate (EUR) in the pre- and post-treatment models, respectively.
LARGE SCALE DATA
N/A
LIMITATIONS, REASONS FOR CAUTION
Our study is limited by the retrospective and single-center design, including moderate sample size, and long observation time span. Methodologically, several features were not present in our data when refitting the McLernon baseline models.
WIDER IMPLICATIONS OF THE FINDINGS
This work illustrates how discrete time survival analysis ML models can be used to generate individualized, multi-cycle IVF prognoses. These ML models can be readily trained on local populations, which facilitates center-specific predictions. Our findings support ongoing efforts to improve patient counselling through data-driven communication about the individualized treatment plans.
FUNDING
This study was supported by grants from the national Research Foundation – Flanders (FWO) (personal mandate 1235924N to FKN, research grant G046024N to CV, TTS) and the Flemish Government (Flanders AI Research program).
DISCLOSURES
The funders had no role in the study design, data collection or analysis, publication decision or manuscript preparation. There are no conflicts of interest to be declared.