DOI: 10.3390/app16199423 ISSN: 2076-3417

AI-Assisted Rehabilitation Assessment for Pediatric Stroke: A Methodological Framework with Explainable Speech Concept Modeling, Video-Based Motor Representation, and Age-Conditioned Normalization

Wangting Liang, Zhiwei Chen, Jiang Xu, Qianqing Li, Jiaxi Wang

This study proposes an explainable framework for developmentally informed pediatric rehabilitation and evaluates audited components on three separate public datasets. Pediatric stroke is a prospective application; no pediatric stroke cohort or synchronized multimodal experiment is included. A reproducibility audit replaced previously reported speech results that used non-independent source variants. In a new calibration-free experiment using 600 Speech Commands recordings from 343 speakers, all variants of each source and all recordings from each speaker remained in one partition. On the speaker-disjoint test set, a 12-feature random forest achieved 0.822 accuracy (95% participant-clustered bootstrap interval 0.774–0.867), whereas four calculated acoustic proxies and a four-concept Sequential Concept Bottleneck Model each achieved 0.578 accuracy. Performance generally approached chance when an entire perturbation family was withheld, demonstrating transformation specificity rather than clinical speech validity. The audited motor video experiment was view-specific rather than cross-view; its pooled accuracy was 0.202, below the 0.217 majority-class baseline, with zero recall for five rare classes. A participant-grouped random forest using six pose-derived proxies reached 0.184 accuracy and 0.095 macro-F1. In all 49 available BJCMC transcripts, leave-one-child-out age normalization reduced the Spearman correlation between speech rate and age from 0.328 (95% interval 0.065–0.535) to 0.042 (−0.052–0.124). A feature-level attribution diagnostic added repeated random controls and stability analysis. These results support a reproducible component-level proof of concept while identifying substantial requirements for clinical, multimodal, and motor validation.