DOI: 10.3390/bdcc10100323 ISSN: 2504-2289

BiMLI: Bidirectional Link–Modality Interaction for Link Prediction in Attribute-Based Multimodal Knowledge Graphs

Hui Zhang, Qilong Han, Hui Sun, Hongtao Song, Chao Liu

In attribute-based multimodal knowledge graphs, images and text serve as auxiliary entity attributes, providing semantic evidence beyond graph topology for link prediction. However, bidirectional conditional dependencies between an entity’s multimodal representations and its local links—the one-hop factual triples formed by relations and neighboring entities—remain insufficiently modeled: different local links rely on structural, visual, and textual evidence to varying degrees, and conversely, the relevance of each modality differs across local links. Existing methods have progressed from modality alignment and fusion to relation-aware, structure-aware, and adaptive selection; yet they mainly let information flow from links to modalities, while conditioning on each modality to distinguish and weight individual surrounding links remains underexplored, leaving the two selection directions disconnected. We therefore propose BiMLI, a bidirectional link–modality interaction framework. Link→Modality selects relevant structural, visual, and textual evidence conditioned on each local link, whereas Modality→Link dynamically weights an entity’s surrounding links conditioned on each modality representation. Interaction-aware feature fusion and neighborhood aggregation integrate both representations, with cross-modal contrastive regularization as an auxiliary constraint. Under a controlled comparison using ten matched random seeds, BiMLI attains mean MRR scores of 41.39% and 44.28% on DB15K and MKG-W, exceeding ROAD by 1.79 and 6.18 percentage points, respectively. Ablation and fine-grained analyses show that the two directions are functionally asymmetric yet complementary, providing evidence for the value of bidirectional conditional modeling.