DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection
Qing Cheng, Yan Jiang, Yuan Gao, Zeng Gao, Su Liu, Xiaoguang TuIn complex UAV aerial scenes, small object detection tasks face challenges such as extremely low pixel occupancy, strong background interference, and sparse effective features, which are further compounded by environmental factors like low illumination. Consequently, single-modality detection algorithms are prone to severe target feature loss and missed detections. Multi-modal image fusion, which complements the texture details of visible light with the thermal radiation characteristics of infrared, is considered an effective approach to overcome the limitations of single physical imaging. However, conventional fusion mechanisms often suffer from semantic gaps when processing heterogeneous data, easily introducing redundant noise and background false alarms. To further improve the accuracy and robustness of small object detection in UAV aerial scenes, this paper proposes a multi-modal detection network that integrates direction-aware multi-granularity and asymmetric context guidance, termed DMAC-Net. Specifically, a Direction-Aware Granularity Enhancement (DAGE) module is first constructed for unified backbone feature extraction, which captures local directions and contour edges of small objects in UAV aerial images with high sensitivity, and expands the receptive field through a multi-granularity mechanism, effectively suppressing false positives induced by complex backgrounds while enhancing the recall of occluded and weakly featured targets. Additionally, the Asymmetric Context Guided Fusion (ACGF) module builds a spatial mechanism via asymmetric receptive fields and performs semantic soft alignment of cross-modal features with dynamic weight assignment, effectively filtering out artifacts and clutter from cross-modal interaction. Experimental results on multiple aerial datasets, including RGBTDronePerson, AVMS and LLVIP demonstrate that the proposed method outperforms existing mainstream models in terms of overall detection accuracy and missed-detection suppression, while exhibiting strong generalization capability and stability under complex lighting transitions and multi-scale variations in UAV monitoring environments.