DOI: 10.1177/09622802261478572 ISSN: 0962-2802

Bayesian variable selection for joint models of heterogeneous longitudinal variables and a binary outcome

Lingpeng Shan, Michelle J Naughton, Electra D Paskett, Michael L Pennell

Biomedical studies often collect mixed-type longitudinal data (e.g., clinical, lifestyle) to identify predictors of a binary health outcome. A common analytical challenge is to link these irregularly collected variables to the outcome, particularly when the number of potential predictors is large. While joint models (JM) can handle this complex data structure, a critical gap exists, as they lack a formal framework for variable selection, limiting their utility for identifying relevant predictors. This article fills that methodological gap by introducing two novel, structured Bayesian variable selection strategies. Our first approach (JM1) selects predictors for the binary outcome, while our more advanced approach (JM2) performs simultaneous two-level selection for both the outcome and the longitudinal trajectories. Our framework leverages shrinkage priors to handle high-dimensional predictors and interactions, preventing overfitting. To guide final variable selection, we extend false discovery rate-based rules to our complex, multi-part joint model, which also accommodates grouped categorical predictors. We apply our method to Women’s Health Initiative and Life and Longevity After Cancer study data, identifying factors for post-treatment insomnia in breast cancer survivors. Our model identified a key predictor missed by conventional methods. This integrated approach provides a robust, interpretable, and computationally efficient framework, offering substantial gains in statistical power.

More from our Archive