PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception
Minjun Kim, Gyuho ChoiAdvanced driver assistance systems (ADAS) require a comprehensive understanding of multiple tasks related to the physical and mental states of drivers and traffic situations. Existing ADAS studies perform driver emotion recognition (DER), driver behavior recognition (DBR), traffic context recognition (TCR), and vehicle behavior recognition (VBR) using models designed based on single-task learning, thereby failing to reflect the interactions among tasks in real driving environments. This paper proposes perception and recognition with inter-modal selective multi-task learning (PRISM-MTL), an integrated multimodal and multi-task learning framework that jointly recognizes DER, DBR, TCR, and VBR. The proposed PRISM-MTL consists of a hierarchical stage-wise attention network (HSA-Net)-based multimodal encoder that extracts spatial features from heterogeneous multimodal inputs and task-specific modality fusion (TSMF), which selectively learns effective modality information for each task. This design addresses negative transfer, a key challenge in multi-task learning. In the multimodal encoder, HSA-Net extracts visual modality tokens that emphasize global structural patterns and key spatial regions from multi-view images, while Token-SE generates joint modality tokens that reflect the spatial configuration of joint data. TSMF generates task-specific fusion features that selectively emphasize the modality cues for each task. The generated task-specific fusion features are summarized through temporal mean pooling, and final predictions of driver states and traffic situations are produced by each task head. Experimental results show that the proposed PRISM-MTL achieves state-of-the-art performance on the public AIDE database, with an mAcc of 86.25% ± 0.35 for multi-task recognition of driver states and traffic situations.