DOI: 10.3390/app16199685 ISSN: 2076-3417

MSSC-EM: A Multi-Semantic Self-Supervised Collaborative Entity Matching Framework

Yaoli Xu, Chunli Xie, Zhilei Yin, Yongwen Liu

Entity matching (EM) aims to determine whether two tuples from heterogeneous data sources refer to the same real-world entity and serves as a fundamental task in data integration, knowledge graph construction, and intelligent information retrieval. However, existing EM methods still suffer from three major limitations: insufficient semantic representation for heterogeneous tuples, difficulty in constructing reliable supervision signals in zero-shot settings, and inadequate collaboration among multiple semantic views, which together result in limited robustness and unstable performance. To address these limitations, we propose a novel self-supervised entity matching framework based on multi-semantic collaboration (MSSC-EM). MSSC-EM consists of three tightly coupled modules: RAG-based Information Augmentation (RIA), Automatic Data Augmentation (ADA), and Collaborative Learning of Multi-Semantic Features (CL-MSF). In MSSC-EM, first, RIA employs Retrieval-Augmented Generation (RAG) to mine implicit contextual semantics from an external knowledge corpus and enrich the original tuple representations. Second, ADA constructs pseudo-positive and negative labels through an Incremental Knowledge Graph Construction (IKGC) strategy and enhances their reliability via confidence-aware Monte Carlo Dropout (MC-Dropout) filtering, thereby providing high-quality supervision signals without manual annotation. Finally, CL-MSF collaboratively learns digital, structural, and relational semantic features to improve the robustness and discriminative capability of tuple representations. Extensive experimental results on eight real-world EM benchmarks demonstrate that, under the unsupervised evaluation setting, MSSC-EM achieves the best or tied-best performance among the compared unsupervised methods. Under the supervised evaluation setting, MSSC-EM further outperforms or matches the compared self-supervised methods on most datasets and achieves competitive results against the compared supervised baselines.