DOI: 10.1049/ipr2.70441 ISSN: 1751-9659

Bi‐SFNet: Bidirectional Cross‐Spatial Fusion Network With Inverse Projection Enhancement for Point Cloud Semantic Segmentation

Runyue Wang, Zhenglei Dou, Yuwei Chen, Senyuan Wang, Shouzheng Zhu, Xinyuan Zhang, Hongxing Qi

ABSTRACT

With the rapid advancement of autonomous driving and environmental perception technologies, acquiring accurate 3D semantic information from LiDAR point clouds has become increasingly important. Currently, multi‐modal fusion is a mainstream direction for improving 3D semantic segmentation performance. However, existing fusion networks face the challenge of efficiently and precisely integrating global semantic context from images with local geometric details from LiDAR. Specifically, fused Bird's‐Eye View (BEV) features are difficult to effectively and finely “inverse‐project” back to the point‐wise local features of the original 3D point cloud. This results in the inability to fully translate fusion advantages into point‐level segmentation accuracy, especially when identifying rare and long‐tail categories (e.g., motorcyclists) and objects under challenging conditions such as occlusion or extreme distance and distant small objects. To address these issues, we propose the Bidirectional Cross‐Spatial Fusion Network (Bi‐SFNet). This network adopts a multi‐stage progressive fusion strategy: it fuses image and LiDAR features in the BEV space via a BEV‐CMF module, utilising a region‐aware mechanism to dynamically guide the adoption of image semantics. Subsequently, the 3D‐IPE path maps the BEV‐fused global features back to the 3D point cloud space. Finally, a secondary deep fusion is performed in the 3D‐IPE module via an Adaptive Gating Fusion (AGF) mechanism, achieving global semantic‐guided dynamic feature enhancement. Experiments on the large‐scale SemanticKITTI dataset demonstrate that Bi‐SFNet achieves an overall mIoU of 71.9%, outperforming the baseline by 6.0 percentage points. It shows significant advantages in recognising rare dynamic small objects, such as motorcycles (94.9%) and bicyclists (94.3%), providing a robust solution for complex 3D perception.

More from our Archive