An independent generative verification agent for in vitro fertilization outcome prediction: Conditional score-based diffusion modeling to cross-check black-box classifiers
Sergei Sergeev, Iuliia Diakova, Lasha NadirashviliObjective
To evaluate a generative model as an independent verification agent for black-box classifiers that predict the outcome of an in vitro fertilization cycle.
Design
Prediction-model development and validation using a held-out internal test set, two prospective observational cohorts, and two public external datasets.
Subjects
The development database contained 15,193 consecutive IVF/ICSI cycles from three centers; 1,520 cycles were held out before development. Prospective cohorts included 96 and 38 cycles with observed outcomes, and external validation included 489 transfer cycles.
Exposure
No intervention was assigned. From seven variables available after the day-1 fertilization check, a conditional score-based diffusion model generated distributions of total and good-quality blastocyst counts. A calibrated LightGBM head estimated clinical pregnancy, split conformal prediction provided count intervals, and TabPFN served as an architecturally independent black-box comparator.
Main Outcome Measures
Clinical pregnancy, defined as fetal cardiac activity on ultrasonography 25 days after transfer; the risk of no blastocyst or no good-quality blastocyst; discrimination, calibration, probabilistic accuracy, and conformal interval coverage.
Results
On retrospective testing, the verification agent achieved AUROC 0.661 (95% CI, 0.631–0.691), Brier score 0.209, and expected calibration error 0.029; predicted and observed pregnancy rates were 33.9% and 34.0%. Ninety-percent conformal intervals covered 93.2% and 91.0% of total- and good-quality-blastocyst counts. Prospectively, AUROC was 0.637 versus 0.535 for TabPFN in the 96-cycle cohort; in the 38-cycle cohort the full pipeline and TabPFN were indistinguishable (0.726 versus 0.700). Externally, discrimination was modest for all systems (0.569–0.607), and risks were overpredicted. The same generated distributions identified cycles yielding no blastocyst (AUROC 0.840, 95% CI, 0.820–0.859) and no good-quality blastocyst (0.883, 0.865–0.900), while underestimating both absolute risks.
Conclusion
Generative trajectories can make disagreement with a black-box predictor inspectable and can identify cycles at risk of futility. Setting-dependent discrimination and calibration require local evaluation and, where justified, model updating before clinical use.