SDD-PCN: Structure-Aware Dual-Decoder Learning for Joint Completion and Semantic Segmentation of Partial Indoor Point Clouds
Xiangrong Ni, Hanqiang Deng, Hao Chen, Sihang Zhou, Jian HuangPartial indoor point clouds acquired by light detection and ranging (LiDAR) are frequently affected by occlusion, restricted viewpoints, and non-uniform sampling, making it difficult to recover complete scene geometry while preserving the semantics of architectural structures and interior objects. To address this challenge, we propose the Structure-Aware Dual-Decoder Point Completion Network (SDD-PCN) for joint point-cloud completion and semantic segmentation. SDD-PCN first separates architectural structures from interior objects and applies dedicated multi-resolution encoders to capture their distinct geometric characteristics. The resulting local features are fused with a global scene representation and processed by parallel completion and segmentation streams that exchange information during progressive decoding. We further construct Partial-S3DIS, a partial-scene extension of the Stanford Large-Scale 3D Indoor Spaces (S3DIS) dataset, containing 1632 partial–complete scene pairs generated from 272 annotated rooms through sensor-aware LiDAR simulation. Experiments show that SDD-PCN achieves L1 and squared L2 Chamfer distance (CD-L1 and CD-L2) values of 16.73 and 15.27, respectively, under the ×10−3 reporting convention. Relative to ProtoFormer, the strongest overall baseline, SDD-PCN reduces CD-L1 by 3.96% and CD-L2 by 7.90%, while obtaining the best completion performance in eight of ten scene categories. The model also achieves an overall accuracy (OA) of 87.23%, a mean class accuracy (mAcc) of 74.33%, and a mean intersection over union (mIoU) of 63.62%, outperforming both task-matched semantic scene completion baselines across all three semantic metrics. Ablation and robustness analyses confirm that structure-aware encoding and dual-stream feature interaction improve geometric recovery, particularly under moderate and severe missingness. These results demonstrate that explicitly modeling indoor structural heterogeneity provides an effective basis for unified point-based scene completion and semantic understanding.