DOI: 10.3390/jimaging12080370 ISSN: 2313-433X

Open Long-Tailed Multimodal 3D Model Classification Based on Sample-Enhanced Category-Space Learning

Yuansa Wang, Xueyao Gao, Chunxiang Zhang, Yongzeng Xue

With the rapid development of three-dimensional (3D) sensing technologies, multimodal 3D model classification has achieved significant progress. However, most existing methods are developed under closed and balanced assumptions, which limits their applicability to open long-tailed scenarios with scarce tail classes, ambiguous hard samples, and continuously emerging categories. In this work, we propose sample-enhanced category-space learning (SE-CSL) for open long-tailed multimodal 3D model classification. The proposed method first uses dual-branch modality encoders to extract point-cloud structural representations and multi-view semantic representations. Mamba is then introduced to model global dependencies across heterogeneous modalities and generate a unified global category representation. To improve the robustness of category representation, we design a category-space learning strategy that jointly integrates long-tailed learning, few-shot representation stabilization, and hard-sample enhancement. A long-tail balanced loss, a few-shot stabilization loss, and a hard-sample boundary loss are further developed to optimize intra-class compactness, inter-class separability, and boundary discrimination. To handle continually emerging classes, we introduce an incremental category-space expansion mechanism that distinguishes new classes, preserves old-class information, and supports unified classification of old and new categories. Extensive experiments on ModelNet40 and ShapeNet55 demonstrate the effectiveness and robustness of SE-CSL.

More from our Archive