Cue-Grounded Flight-Intent Understanding from Cinematic Air-Combat Scenes via Noise-Robust Multimodal Perception and Fusion
Zhiyuan Xiang, Bo Ma, Shu Chen, Bin Chen, Xiang Deng, Yicheng Su, Jie RenFlight intent cannot always be inferred from aircraft motion because the purpose of a maneuver depends on its surrounding context. This study examines multimodal flight-intent understanding in cinematic air-combat footage. We introduce CineAeroIntent, a benchmark constructed from 3972 annotated candidate clips collected from Chinese air-combat films and television programs. After quality control, 3958 clips were retained and organized by observation perspective. Each sample contains an intent label and an explanation linked to evidence in the clip. Using Qwen2.5-Omni as the backbone, we evaluate Modality-Decoupled Expert Projection (MDEP), which assigns a separate low-rank update to each input stream, and Post-Projection Value-Gating (PPVG), which reduces the contribution of low-utility value representations before attention aggregation. The full model obtains 89.6% overall accuracy and 88.4% Macro F1. Under matched adaptation settings, accuracy increases from 82.3% with shared LoRA to 86.1% with MDEP and 89.6% with MDEP–PPVG. In a natural-subset stress test, the full model retains 95.7% of its lower-interference accuracy on clips with prominent background music, compared with 82.6% for the unadapted backbone. These results support cue-grounded multimodal modeling for cinematic flight-scene understanding.