Performance of deep learning auto-segmentation models for cardiac substructures in paediatric proton beam therapy
L Goyal, R Bell, K Banfill, S Ingram, M Lowe, H Mandeville, E Osorio, M Shen, S Pan, Y Wang, E Smith, M AznarAbstract
Background
Accurate segmentation of cardiac chambers and substructures is fundamental for quantifying radiation exposure and studying late cardiovascular side-effects in children receiving thoracic radiotherapy. Manual segmentation is time consuming and subject to inter-observer variability, therefore limiting its use in cardio-oncology research. Automated approaches may facilitate reproducible cardiac substructure dose assessment; however, most available models are trained on adult anatomy and remain insufficiently validated in paediatric populations.
Methods
Paediatric and young adult patients treated with thoracic proton beam therapy (PBT) (an advanced form of radiotherapy) at a single UK centre were retrospectively identified. Cardiac substructures were manually segmented on pre-treatment CT scans by a clinical oncologist using established cardiac segmentation atlases [1-3]. Two automated models were tested: 1) PlatipPy, an open-source hybrid framework combining deep learning based whole heart segmentation with anatomically constrained substructure delineation [4], and 2) Limbus AI (version 1.9.0-B1), a commercial deep learning segmentation model [5]. Geometric agreement between automated and manual segmentations was evaluated using the Dice similarity coefficient (DSC, a measure of overlap between structures, ideal value=1), and mean distance to agreement (MDA, a measure of distance between surfaces, ideal value=0).
Results
Twenty-five patients were included (median age 11 years, range 3–18; 18 male). Diagnoses included paediatric Sarcoma (n=18), lymphoma (n=5), neuroblastoma (n=2). Median (range) prescription dose was 50.2 Gy (10–55.8) in 17 (10–31) fractions.
Whole-heart segmentation showed high agreement for both models, with median DSC of 0.88 for PlatipPy and 0.88 for Limbus AI, and low median MDA values 3.02 mm and 3.20 mm, respectively.
Performance was reduced for individual chambers, particularly atria, with median DSC of 0.59–0.61 for PlatipPy and 0.61–0.65 for Limbus AI, and median MDA of approximately 4.72–5.81 mm. Ventricular structures showed better agreement (DSC 0.69–0.76 and 0.65–0.78; MDA <4.6 mm). Markedly poorer performance was observed for smaller substructures, including coronary arteries, valves and conduction tissue, with median DSC generally below 0.35 and large spatial discrepancies. Limbus AI failed to generate contours for several small structures, whereas PlatipPy consistently produced segmentations. Detailed results are summarised in Tables 1 and 2.
Conclusion
Both models performed adequately for the whole heart, but led to suboptimal results for cardiac substructures, poorer than what is observed in reports of auto segmentation in adult patients [1]. This emphasises the need for deep learning models specifically trained on paediatric data. Further work will include an assessment of the dosimetric impact of segmentation differences on cardiac substructure dose estimates in children treated with PBT.