DOI: 10.1142/s1469026826500367 ISSN: 1469-0268

From Noise to Knowledge: Distill Social Media Opinions on Rumor Detection Exploiting LLMs

Shakshi Sharma, Anjali Goyal, Naman Ahuja, Bibhudutta Pati, Rajesh Sharma, Ponnurangam Kumaraguru

The majority of social media research relies on datasets gathered from social media websites such as Reddit and Twitter. However, their intrinsic high noisy content reduces the performance of data-driven models. This restriction of noisy data has made data preparation techniques necessary. Current systems usually need to pay more attention to data quality and their performance heavily depends on noisy social media datasets. In this work, we concentrate on creating high-quality social media data for rumor detection tasks on the widely popular PHEME-9 dataset. Eliminating noisy responses is essential for misinformation tasks since it guarantees that the model has been trained on precise and pertinent data, improving its capacity to identify and validate rumors successfully. Large language models (LLMs) are used in this work to filter out irrelevant comments prior to the application of machine and deep learning techniques. This bifurcated approach helps to improve model accuracy and lower computing burden. We assume that this will further help in accurate rumor identification and can support environmental sustainability using less computational resources. Our proposed methodology shows an average improvement in the trained filtered models’ performance in terms of accuracy and F1-scores on six events in the PHEME-9 dataset. Further, to validate the effectiveness of the trained model, we performed interpretability and error analysis.

More from our Archive