MOT-Assisted Object-Level Point Cloud Extraction from Multi-View Observations
Cheng Ju, Zejing Zhao, Chaowen Shen, Xuantong Li, Akio NamikiObject-level point cloud extraction is relevant to robotic perception and may provide useful object-centric observations for mapping and SLAM. Many existing methods rely on per-frame instance segmentation, resulting in high computational cost and limited ability to extract multiple objects simultaneously in multi-object scenarios. This paper proposes an efficient MOT-assisted pipeline for multi-view object point cloud extraction. The pipeline first employs 2D multi-object tracking (MOT) to establish consistent object correspondences across views, and then combines monocular depth-based reconstruction with multi-view geometric association to estimate coarse object locations. An adaptive spherical proposal and a density-based refinement strategy are further introduced to extract clean object-specific point clouds while suppressing background noise and outliers. Experiments on the DTU, MVImgNet, and ScanNet++ datasets demonstrate the effectiveness of the proposed method. Relative to the unprocessed scene-level point cloud, the Chamfer Distance is reduced from 90.97 mm to 12.52 mm and Precision increases from 0.2998 to 0.8776 on DTU dataset Sequence 30. On MVImgNet, the Chamfer Distance decreases from 1.23 m to 0.46 m and Precision increases from 0.4777 to 0.9337, demonstrating effective removal of non-target scene points under real-world viewing conditions. Moreover, the proposed method reduces the cost of object-level association and provides competitive object-level extraction quality under the tested settings. Compared with the closely matched segmentation-based extraction, the proposed method provides lower measured object extraction runtime and built-in cross-view Track-ID association, while sacrificing some boundary accuracy.