DOI: 10.3390/electronics15194474 ISSN: 2079-9292

MKE: A Lightweight Framework for Multimodal Knowledge Extraction via Knowledge Distillation

Shixiong Liu, Xuming Ye, Xingyu Wang, Lianshuai Wang, Weiyu Dong, Meng An

Multimodal knowledge extraction has become increasingly important for applications that require compact and informative representations of text-image data, yet existing models often remain computationally expensive for practical use. To address this issue, we propose MKE, a lightweight framework for multimodal knowledge extraction based on knowledge distillation and efficient multimodal fusion. Specifically, MKE transfers linguistic knowledge from a large BART teacher model to a compact StuBART backbone and combines Flash-Attention-based cross-modal interaction with a lightweight nonlinear transformation module. Experiments on the MSMO English news dataset show that MKE achieves competitive multimodal summarization performance with 270M parameters and a computational cost of 155.63 GMACs (311.27 GFLOPs) per sample. We further evaluate inference efficiency under a controlled RTX 4090 setup, and we clarify that deployment on smartphones, IoT systems, and wearable devices remains a promising direction for future work rather than a scenario directly validated in this study.