Machine-Learning-Based Prediction of Cervical Pedicle Screw Malposition from Clinical and Anatomical Features
Milan S. Vosko, Stefan Aspalter, Anja Blenk, Petra Böhm, Nico Stroh-Holly, Andreas Gruber, Wolfgang SenkerBackground/Objectives: Cervical pedicle screw (CPS) placement provides superior biomechanical stability but remains technically demanding and associated with a risk of screw malposition. While recent advances in imaging and navigation have improved placement accuracy, reliable prediction of malposition remains challenging. The aim of this study was to evaluate whether machine learning (ML) models can predict CPS malposition using structured clinical and anatomical features. Methods: We performed a retrospective analysis of 862 pedicle screws from 168 posterior cervical spine surgeries conducted at our institution between 2018 and 2025. Clinical, procedural, and anatomical variables, including age, sex, body size parameters, surgical indication, vertebral level, pedicle angle, and pedicle width, were evaluated. Pedicle morphology was partially derived from CT-based automated segmentation using TotalSegmentator (v2.13.0), while selected anatomical parameters were manually measured. Supervised ML models, including Random Forest, Balanced Random Forest, XGBoost (v3.2.0), Support Vector Machine, and K-Nearest Neighbor, were trained and compared using Python and scikit-learn to predict inaccurate screw placement. Model performance was evaluated using Area Under the Receiver Operating Characteristic Curve (ROC AUC), F1-score, precision, and recall. Model interpretability was assessed using Shapley Additive Explanations (SHAP). Results: The dataset showed a clinically representative class distribution, with 91.1% of screws classified as acceptable and 8.9% as inaccurate. Across all models, predictive performance was moderate and consistent. Balanced Random Forest achieved the highest discriminative performance (ROC AUC 0.69) and provided the most balanced classification profile, while other models demonstrated comparable overall performance with varying sensitivity to the minority class. SHAP analysis identified anatomical and procedural variables, including pedicle width and angle, as relevant contributors to model output. Feature contributions were distributed across variables, with substantial overlap between outcome groups. Conclusions: ML-based prediction of CPS malposition using clinical and anatomical features demonstrates consistent and interpretable performance. The results highlight that predictive performance is primarily influenced by dataset characteristics, including class distribution and feature overlap, rather than model selection alone. This study provides an important baseline for ML-based CPS prediction and supports future research integrating larger datasets and more detailed anatomical representations to enhance predictive accuracy.