Auditing Multimodal Fusion Benchmarks: Leakage, Attribution, and Calibration in Elderly Fall Detection
Abha Tewari, Dhananjay KalbandeReported gains from multimodal fusion, feature-attribution explanations, and model-confidence scores are only as trustworthy as the evaluation protocol and preprocessing pipeline behind them, yet these are rarely audited together. We develop a benchmarking-and-auditing methodology addressing four failure points jointly, subject-overlapping evaluation, class-imbalance correction under a subject-wise split, feature attribution, and calibration, instantiated on seven models under leave-one-subject-out (LOSO) cross-validation spanning three encoder-backbone pairings and early, late, and hybrid fusion, using VARISHTA-MM50, a synchronised IMU-and-video dataset of 50 elderly participants (aged 60–90). Random Forest under late fusion achieves the best LOSO accuracy (85.00%), outperforming every deep architecture tested. Subject-overlapping evaluation overstates accuracy by a mean of 8.32 percentage points; a GroupKFold control isolates subject overlap as the cause, and the effect replicates on an independent dataset. A leakage-safe oversampling protocol confirms the model ranking is not a preprocessing artefact. TreeSHAP attribution shows individual inertial features outweigh individual video features (1.32× per feature), though the wider video branch contributes more in aggregate, and a modality-removal ablation shows this dependence is specific to the tree ensemble. The most accurate model is the least well-calibrated by Expected Calibration Error. Subject-independent benchmarking, leakage auditing, attribution, and calibration form one connected evaluation discipline rather than four separate checks.