DOI: 10.3390/robotics15090175 ISSN: 2218-6581

Open-Vocabulary Instance Segmentation for Scene Understanding in Mobile Robots

Macoris Decena-Gimenez, Pepe Ojeda, Jose-Raul Ruiz-Sarmiento, Javier Gonzalez-Jimenez

High-level robotic tasks, such as those involving planning and interaction, demand a certain degree of scene understanding through suitable representations of the environment that enrich geometric information with object-level semantics, commonly referred to as semantic maps. Traditional techniques to build these maps are restricted to working with a set of categories that is fixed during training, which limits the information the resulting maps can contain. To overcome this limitation, this work presents an open-vocabulary semantic mapping pipeline that integrates TALOS (TAgging–LOcation–Segmentation) with the probabilistic, instance-aware Voxeland framework. TALOS employs a modular generative approach in which vision–language and large language models infer categories for the objects in the scene, an open-vocabulary grounding model localizes their instances, and a category-agnostic segmentation model produces pixel-accurate masks. These predictions are combined with depth data and fused into persistent 3D maps using Voxeland while preserving geometric, semantic, and instance uncertainty. The approach was evaluated on diverse ScanNet v2 scenes and compared with two alternative perception front-ends: the closed-vocabulary Detectron2/Mask R-CNN and the open-vocabulary discriminative model YOLOE. Across five independent runs, TALOS achieves a mean mAP@0.5 of 0.1607±0.0255, compared with 0.0834 for YOLOE and 0.0453 for Detectron2/Mask R-CNN, and exceeds both baselines in every run. Qualitative results further show a better balance between map completeness, geometric clarity, semantic coherence, and instance separation.