MindVoice: A Web-Based Multi-Dataset Speech Depression Screening System with Pipeline-Aware Grad-CAM Explainability
Arjita Choubey, Manoj Kumar Pandey, Ashwani Kumar Dubey, Alvaro RochaAlthough speech-based depression assessment has received considerable attention, achieving reliable confidence and cross-corpus applicability with explanation of prediction still remains challenging. This study presents a reliability-aware framework for speech-only depression severity assessment and its deployment through MindVoice. The distress analysis interview corpus (DAIC), extended distress analysis interview corpus (E-DAIC) and multimodal open dataset for mental disorder analysis (MODMA) are integrated to form a heterogeneous multi-corpus cross-lingual dataset of three classes. The Audio-Textual Depression Corpus (EATD) is also evaluated for four classes. Domain-specific feature normalization is applied to reduce variations arising from differences among the datasets. CNN-derived speech representations are subsequently refined through Analysis of Variance (ANOVA) followed by artificial bee colony (ABC)-based feature optimization, and an MLP assigns samples to severity levels. To improve the trustworthiness of predicted probabilities, seven calibration approaches like Temperature, Platt (Sigmoid), Isotonic, Vector, Matrix, Dirichlet, and Beta calibration are compared using accuracy and Expected calibration error. Reliability is quantified using a fixed-weight combination of bin-based reliability (0.7) and entropy (0.3) while comparing with other reliability assessment methods. A pipeline-aware gradient-weighted class activation mapping approach is proposed that enables explanations by reconstructing sklearn components as Tensorflow components for gradient backpropagation across the hybrid pipeline. The framework obtains 76.5% accuracy on the combined corpus, while the DAIC, E-DAIC, and MODMA achieve an accuracy of 81.4%, 72% and 80% respectively, and the EATD achieves 78.5% accuracy for 4-class classification. The proposed framework provides effective classification for a heterogeneous dataset. Its deployment on the web as the MindVoice application inculcates reliability-aware and explainable prediction that provides a more trustworthy system than the conventional performance-only assessment.