DOI: 10.1145/3831951 ISSN: 2474-9567

LEMuR: Language Embedded 3D Segmentation and Object Tracking for Mixed Reality

Wendi Yu, Zhongyi Bai, Hongyu Zhou, Brandon Victor Syiem, Eduardo Velloso

Mixed reality enables distributed collaboration across distinct physical spaces. However, collaborator interactions are restricted from shared manipulation of physical objects in their surroundings to aid communication. Existing approaches either substitute low-fidelity virtual proxies for physical objects, stripping away the communicative context essential for interpreting actions, or rely on external cameras that constrains interaction to fixed capture regions and impede practical deployment. We introduce LEMuR, a language embedded 3D segmentation and tracking pipeline that, following an offline static scene reconstruction phase, enables real-time sharing and synchronization of everyday physical objects in MR. Users specify objects to share through natural language queries, then LEMuR selectively visualizes it and synchronizes local manipulations as high-fidelity reconstructions for remote collaborators. Through a user study with 12 participants grounded in Knapp's framework of nonverbal communication, we examined how LEMuR supports object-mediated gestures, proxemics, and artifact cues, revealing emergent interaction patterns unique to hybrid physical-digital collaboration. We further evaluate pipeline performance and discuss design implications for object-centred multi-user MR systems.