MSSC-EM: A Multi-Semantic Self-Supervised Collaborative Entity Matching Framework
Yaoli Xu, Chunli Xie, Zhilei Yin, Yongwen LiuEntity matching (EM) aims to determine whether two tuples from heterogeneous data sources refer to the same real-world entity and serves as a fundamental task in data integration, knowledge graph construction, and intelligent information retrieval. However, existing EM methods still suffer from three major limitations: insufficient semantic representation for heterogeneous tuples, difficulty in constructing reliable supervision signals in zero-shot settings, and inadequate collaboration among multiple semantic views, which together result in limited robustness and unstable performance. To address these limitations, we propose a novel self-supervised entity matching framework based on multi-semantic collaboration (MSSC-EM). MSSC-EM consists of three tightly coupled modules: RAG-based Information Augmentation (RIA), Automatic Data Augmentation (ADA), and Collaborative Learning of Multi-Semantic Features (CL-MSF). In MSSC-EM, first, RIA employs Retrieval-Augmented Generation (RAG) to mine implicit contextual semantics from an external knowledge corpus and enrich the original tuple representations. Second, ADA constructs pseudo-positive and negative labels through an Incremental Knowledge Graph Construction (IKGC) strategy and enhances their reliability via confidence-aware Monte Carlo Dropout (MC-Dropout) filtering, thereby providing high-quality supervision signals without manual annotation. Finally, CL-MSF collaboratively learns digital, structural, and relational semantic features to improve the robustness and discriminative capability of tuple representations. Extensive experimental results on eight real-world EM benchmarks demonstrate that, under the unsupervised evaluation setting, MSSC-EM achieves the best or tied-best performance among the compared unsupervised methods. Under the supervised evaluation setting, MSSC-EM further outperforms or matches the compared self-supervised methods on most datasets and achieves competitive results against the compared supervised baselines.