Synergistic Lightweight Bilinear Pooling Fusion for Fake News Detection
Jun Li, Jun Li, Linghao Yan, Junnan JiangTo improve cross-modal interaction modeling while controlling the complexity of multimodal fusion, this paper proposes a multimodal fake news detection framework that integrates textual and visual information through cross-modal attention and a lightweight bilinear pooling module at the fusion stage, while retaining BERT-BiLSTM and ResNet-50 as the textual and visual encoders. Cross-modal attention performs text-conditioned visual weighting, while lightweight bilinear pooling captures compact second-order cross-modal interactions. The lightweight bilinear pooling module reduces the dimensional and parameter overhead associated with conventional full bilinear interaction while preserving effective cross-modal interaction modeling. The model achieves an accuracy of nearly 93% on the Weibo dataset, outperforming the compared baseline methods. Ablation studies and projection-dimension analyses further support the effectiveness and complementary roles of the two fusion components under the evaluated Weibo setting.