DOI: 10.3390/sym18101580 ISSN: 2073-8994

A Cross-Modal Consistency and Reliability-Aware Fusion Network for Railway Foreign-Object Segmentation

Qian Wang, Peng Dai, Zizhen Xu, Haoran Song, Jing Shi, Lingchong Wang, Zhi Wang, Dongxu Hong

Real-time segmentation of foreign objects between rails is essential for high-speed railway inspection. Single-modality perception remains limited because two-dimensional images provide rich texture but lack depth, whereas point clouds provide geometry but are weak in color and material discrimination. To address this problem, we developed a high-speed RGB-depth acquisition workflow and propose a cross-modal consistency and reliability-aware fusion network for railway foreign-object segmentation. The network uses a dual-branch image-point cloud architecture and a Dual-Input Attention (DIA) module to align hierarchical features and adaptively weight geometric and texture cues after converting local uncertainty into reliability scores. Experiments on 885 image-point cloud pairs collected from an in-service railway inspection train operating at approximately 100 km/h show that the proposed model achieves 96.210% Recall, 96.952% OA and 94.893% IoU. Compared with the strongest point-cloud baseline in this study, the model improves IoU by 0.703 percentage points while using fewer FLOPs than Point Transformer v3. The numerical comparison is interpreted together with module, fusion-depth, sampling-ratio, loss-weight, input-modality and density ablation studies, and the remaining need for additional multimodal baselines and multi-seed validation is now stated explicitly.