Developing an integrated algorithm to support autism diagnostic decisions: Model performance and statistical fairness
Aaron J. Kaat, Ashlynn Campagna, Hannah Feiner, Jeffrey Michael Grauzer, Amanda Nicole Hobson, Maranda Jones, Jordan Paige Lee, Laura Jean Sudec, Megan York RobertsAbstract
Background
Best practices in diagnosing autism spectrum disorder require an expert diagnostician to integrate multiple sources of information, including direct observations, interviews, and rating scales completed by knowledgeable informants. Machine learning methods are well‐positioned to integrate disparate information sources to improve classification. This study sought to evaluate whether a machine learning algorithm could support appropriately weighting multiple sources of diagnostically relevant information.
Methods
Telehealth diagnostic assessments were conducted for 639 toddlers (age mean = 30.4, SD = 4.3 months; 29.7% female) already receiving early intervention (EI). The Toddler Autism Symptom Interview (TASI) and the TELE‐ASD‐PEDS (TAP) were completed during separate visits. Each child's caregivers and usual EI providers asynchronously completed rating scales of autistic characteristics developed for this study. Calibration and validation data subsets were created using a 2:1 split. The optimal algorithm was developed using elastic net regularized regression in the calibration dataset, which was then evaluated in the validation dataset. Statistical fairness criteria evaluated whether the proposed algorithm functioned similarly across multiple protected group statuses.
Results
The prevalence of autism in the sample was 80.4%. The algorithm developed in the calibration dataset included both the caregiver‐ and EI provider‐completed rating scales, four items from the TASI, and six items from the TAP. Model performance was high in both the calibration and validation samples (sensitivity >0.90; specificity >0.75; kappa >0.60), but statistical fairness criteria varied. False positives were relatively rare, but were more common in advantaged groups, which may suggest systematic under‐classification by the algorithm in minoritized groups.
Conclusion
Diagnostic practices for autism require integrating multiple sources of information. Machine learning can support diagnosticians in this effort, as demonstrated here utilizing multiple rating scales, the TASI, and the TAP. However, there is a risk for bias when using machine learning and as such, no algorithm should replace expert clinical judgment.