DOI: 10.3390/s26165052 ISSN: 1424-8220

PartSense-IP: Part-Aware Vision–Language Sensor Fusion for Visual–Semantic Consistency Evaluation of IP Prototypes

Yangfan Feng, Wen Zhao

Evaluating whether an intellectual property (IP) prototype faithfully preserves the visual identity and semantic intent of its original concept design is an important yet challenging task in product design and creative prototyping. Existing evaluation practices mainly rely on manual inspection or global image-level similarity comparison, which are subjective, difficult to reproduce, and insufficient for localizing identity-critical deviations. To address this problem, this paper proposes PartSense-IP, a part-aware vision–language sensor fusion framework for visual–semantic consistency evaluation of IP prototypes. The proposed framework takes a 2D concept image, an optional textual design description, and multi-view RGB-D sensor observations of a prototype as inputs. It first constructs a multi-view prototype representation and decomposes both the concept and prototype observations into design-relevant parts. Dense visual features, color and shape descriptors, and vision–language semantic embeddings are then extracted to evaluate part-level consistency. A Part-Aware Visual–Semantic Consistency Fusion (PVCF) algorithm is further developed to integrate shape, color, local visual similarity, semantic alignment, and cross-view stability into a unified IP consistency score. In addition to scalar scoring, PartSense-IP generates localized difference maps, 3D inconsistency visualization, and interpretable design feedback for prototype refinement. Experiments on the proposed IP-ProtoSense evaluation protocol demonstrate that PartSense-IP outperforms representative vision–language, dense-visual, segmentation-based, and 3D multimodal baselines in consistency scoring, inconsistency detection, localization, ablation, and robustness evaluation.

More from our Archive