Privacy‐Preserving Federated Audio‐Visual Saliency Analysis and Data Sharing Architecture for Rehabilitation of People With Disabilities
Ning FangABSTRACT
Privacy constraints on facial, movement, and voice data restrict centralized model training across rehabilitation institutions. This study presents a privacy‐preserving federated audio‐visual saliency architecture, termed PPFAV, for edge‐based rehabilitation assessment. A three‐stream encoder represents speech instructions, dynamic action sequences, and static visual cues, while a self‐supervised cross‐modal alignment objective links synchronized instructions and movement responses. Before aggregation, clipped local updates are perturbed with calibrated Laplace noise so that raw samples remain at their local clients. Against the movement‐derived pseudo‐saliency references under the reported Non‐IID simulation, PPFAV achieved a correlation coefficient of 0.82 and a similarity score of 0.76, while retaining an F1‐score of 0.81 at = 0.1. These results indicate that cross‐modal alignment and privacy‐aware federated aggregation can support robust attention localization under simulated client heterogeneity. However, the evaluation is limited to a small controlled rehabilitation dataset with synthesized audio and simulated clients and networks; clinical validity, saliency‐ground‐truth reliability, and cumulative privacy accounting therefore require further verification.