DOI: 10.1145/3849805 ISSN: 1046-8188
Filtering LLM-Synthesized Data for Cold-Start Recommendation: A Progressive Influence Function Approach
Tianhao Shi, Yang Zhang, Xiaoxue Zang, Chenyi Lei, Han Li, Yang Song, Fuli Feng Leveraging Large Language Models (LLMs)-synthesized data to enable representation learning for cold-start entities has emerged as a promising solution for addressing cold-start recommendations. However, existing studies often over-rely on LLM-generated data while neglecting the detrimental noise it may introduce. This work investigates data filtering mechanisms to remove detrimental synthetic samples.
One promising approach is to use
influence functions
to measure the influence of each data. Nevertheless, a technical challenge arises because accurate influence estimation requires an unbiased reference model, which cannot be satisfied due to the existence of unknown noise. To address this, we conduct a theoretical analysis, revealing that when the goal is to identify a small subset of highly harmful data, even a biased model can still reliably capture the most detrimental points. Building on this insight, we propose a
Progressive Influence Function-based method
(PIF) for data filtering. PIF iteratively refines influence estimates: it repeatedly trains a new reference model after removing the previous step's most harmful data and reassesses influence. As detrimental data is progressively removed, the reference model becomes less biased. Finally, PIF ensembles influence estimates across iterations for robust filtering. Extensive experiments validate the efficacy of our approach.