DOI: 10.1145/3848635 ISSN: 2157-6904

IMoKGNN: Dual-Stream Fusion of Generic and Task-Specific Language Model Features for Graph Neural Networks

Hao Yan, Chaozhuo Li, Jun Yin, Weihao Han, Hao Sun, Senzhang Wang, Jian Zhang, Jianxin Wang

Text-Attributed Graphs (TAGs) are prevalent in various real-world scenarios, where each node is associated with a text attribute. Representation learning on TAGs relies on a comprehensive understanding of both the textual attributes and the topological connections. Recent works have enhanced graph neural networks (GNNs) with pre-trained language models (PLMs) for textual attribute modeling, achieving promising results compared with shallow text representations such as BoW. With the advent of more powerful LLMs, how to effectively and efficiently inject the generic knowledge in LLMs into representation learning on TAGs still calls for systematic investigation. To this end, we first examine whether mainstream generative LLMs can provide high-quality generic text embeddings that can be directly used by downstream GNNs, and further explore their complementarity with the task-specific knowledge from a tuned PLM (TLM). Based on these observations, we propose a dual-channel knowledge fusion framework, named MoKGNN, to simultaneously leverage the generic semantic knowledge from LLMs and the task adaptation capability of TLM. Specifically, we design a dual-channel feature extraction module to learn generic and task-specific text representations separately. Next, we introduce a Knowledge Alignment Network to align the LLM embeddings with the TLM representations. Finally, MoKGNN adopts a global mixing weight to fuse the two channels, obtaining a unified node embedding for downstream GNN training. Moreover, we further analyze a key limitation of such early feature fusion under message passing: once heterogeneous text features are merged before neighborhood aggregation, source-specific noise and non-aligned semantics may be diffused together and become harder to disentangle and correct. Motivated by this observation, we propose IMoKGNN, a decoupled dual-stream framework that propagates LLM and TLM features in parallel with a shared GNN backbone and performs late fusion in logit space. Extensive experiments on five TAG datasets demonstrate the effectiveness of our approaches for both node classification and link prediction, and show that IMoKGNN yields more reliable improvements than MoKGNN.