Condition-Specific LC Retention Time Prediction: Feature-Selected QSRR Versus Pretrained Graph Isomorphism Network Transfer Learning
Roman Szucs, Emília Sýkorová, Iveta Boháčová, Ilaria Neri, Lucy Morgan, Merisa Moriarty, Melissa Hanna-BrownCondition-specific liquid chromatographic retention time prediction remains challenging because retention depends on both molecular structure and experimental conditions. This study compared feature-selected quantitative structure–retention relationship (FS-QSRR) models with pretrained graph isomorphism network (GIN) transfer learning for three reversed-phase LC datasets measured under acidic, neutral, and basic conditions. Mordred descriptor-based QSRR models were developed using leakage-safe preprocessing, nested cross-validation, and feature selection. The selected FS-QSRR workflow for each dataset was then compared with pretrained GIN transfer learning using identical 100 repeated random 80/20 train–test splits. Feature selection substantially reduced descriptor dimensionality but did not consistently improve predictive accuracy over the best baseline descriptor models. Under matched validation, GIN transfer learning gave lower RMSE for all three datasets, decreasing error from 0.816 to 0.613 min under acidic conditions, from 0.868 to 0.665 min under neutral conditions, and from 1.019 to 0.876 min under basic conditions. The corresponding RMSE reductions were 24.8%, 23.4%, and 14.0%, respectively. Matched prediction error analysis showed that GIN particularly reduced the frequency and magnitude of large errors, although the improvement was more modest under basic conditions. Descriptor frequency analysis revealed condition-dependent contributions from lipophilicity, ionization-related, electronic, and topological descriptor groups. These findings support pretrained GIN transfer learning as the stronger predictive approach, while FS-QSRR remains valuable for model simplification and chemical interpretation.