Multi-Atlas–Based Segmentation of Pediatric Vocal Tract Anatomy in Dynamic Magnetic Resonance Imaging
Hahn Kang, Fangxu Xing, Imani R. Gilbert, Jiyoon Kim, Abigail Kostolansky, Bradley Sutton, Jonghye Woo, Jamie L. PerryPurpose:
Segmentation of velopharyngeal and vocal tract anatomy is crucial for quantifying structures and analyzing dynamics during speech. Automatic methods are needed to replace labor-intensive and poorly reproducible manual segmentation. To address this, we compared four multi-atlas–based methods to determine which provides the most accurate segmentation of the tongue, velum, and adenoid and how segmentation accuracy varies with the number of temporal frames.
Method:
Five spatiotemporal atlases of speech tasks were built, each with around 80 frames. We performed diffeomorphic registration of 10–40 frames of each atlas, then propagated labels to evaluate how the number of frames impacts performance. Four label fusion techniques—majority voting label fusion (MVLF), simultaneous truth and performance level estimation (STAPLE), multi-atlas label fusion (MALF), and corrective learning (CL)—were applied to evaluate segmentation accuracy. Performance was assessed using dice similarity coefficient (DSC) values across increasing numbers of input frames (10–40).
Results:
CL consistently produced the highest segmentation accuracy, with average DSC improving from 0.89 at 10 frames to 0.92 at 40 frames. In contrast, MVLF, STAPLE, and MALF showed relatively flat trends, with DSC values ranging from 0.82 to 0.87. Statistical analysis confirmed that only CL exhibited a significant positive association between frame count and DSC.
Conclusions:
CL demonstrates superior performance in segmenting pediatric vocal tract structures in dynamic magnetic resonance imaging, particularly as the number of temporal frames increases. These findings support the use of CL in analyzing speech-related anatomy and highlight its potential to replace manual segmentation in small specialized data sets.
Supplemental Material: