DOI: 10.1145/3847657 ISSN: 1046-8188

Lightweight Denoising and Aligning for Multi-modal Recommender System

Guipeng Xv, Yi Liu, Xinyu Li, Zhenhua Huang, Chen Lin

Multi-modal recommender system (MRS) has emerged as a key information retrieval technology, widely adopted to enhance various web platforms. However, three interrelated challenges remain insufficiently explored: (1) noisy multi-modal content, (2) noisy user feedback, and (3) misalignment between multi-modal content and user feedback. Previous works have either overlooked these challenges or proposed a complex solution. To tackle these challenges in a lightweight way, we propose L ightweight D enoising and A ligning for M ulti-modal R ecommender S ystem (LDA-MRS). During graph construction, LDA-MRS only constructs a single item-item graph based on consistent cross-modal similarity and dynamic user behavior, effectively reducing noise in multi-modal content. We provide a

Lightweight Static Strategy
and an
Accurate Dynamic Strategy
for fusing the graphs. During supervised learning, LDA-MRS leverages multi-modal content to estimate the probability of pairwise observed feedback and introduces a Lightweight Denoising BPR Loss to effectively denoise user feedback. During alignment, LDA-MRS uses
Lightweight Alignment guided by User preference
to improve task-specific alignment and
Lightweight Alignment guided by graded Item relations
to achieve finer-grained alignment. LDA-MRS is a lightweight and model-agnostic framework. Experiments on different datasets, backbones, and noisy situations show that LDA-MRS consistently delivers significant performance improvements, highlighting the robustness and effectiveness of LDA-MRS in different conditions.