Multi-Domain Spectral and Time-Series Imaging Representations for Pediatric Congenital Heart Disease Classification
Sittinon Thanonklang, Talit Jumphoo, Wongsathon Pathonsuwan, Kasidit Kokkhunthod, Atcharawan Rattanasak, Rattikan Nualsri, Porntip Nimkuntod, Pattama Tongdee, Monthippa Uthansakul, Peerapong UthansakulCongenital heart disease (CHD) is a major cause of infant morbidity and mortality, and timely diagnosis remains difficult where advanced imaging is not consistently available. Automated phonocardiogram (PCG) screening can support early triage, but pediatric multiclass classification is limited by subtle acoustic differences and class imbalance. This study proposes a five-class heart-sound framework (Normal, ASD, PDA, PFO, and VSD) using multi-domain feature fusion and a convolutional recurrent neural network (CRNN). A four-channel representation combining Log-Mel, PCEN-Mel, Gramian Angular Summation Field (GASF), and Markov Transition Field (MTF) is modeled with a CNN encoder, bidirectional GRU, and dual temporal pooling. Recordings are segmented with uniform 50% overlap and evaluated under a strict subject-wise protocol with macro-F1-guided model selection. Across five runs on the pediatric ZCHSound dataset, the five-class model achieves a mean macro F1-score of 78.61 ± 1.00% and a mean per-subject accuracy of 87.94 ± 0.87%. A dedicated Normal-versus-Abnormal model reaches 95.74% accuracy. These results suggest that multi-domain fusion with sequence-aware modeling is a promising approach for pediatric CHD auscultation support.