TDKG: Text-Guided Domain Knowledge Generalization with Cross-Modal Feature Alignment
Silei Shen, Jingyi Zhang, Yuxi Wang, Feng Chen, Yan Liu, Junran Peng, Ziwei ZhuDomain generalization (DG) attempts to generalize a model trained on single or multiple source domains to an unseen target domain. Motivated by the transferability of vision-language pretrained models, we argue that text can provide complementary semantic cues for domain generalization. In this paper, we develop a Text-guided Domain Knowledge Generalization (TDKG) framework with three components. First, we devise an automatic word-generation method that uses lexical substitution to produce domain-relevant descriptors. Second, we embed these descriptors into the text feature space through prompt learning while preserving category semantics and encouraging diversity across domain words. Finally, we use both image and generated text features to train a normalized classifier and update the image encoder. The text branch is required only during training; inference uses the test image alone. Experiments on five domain generalization benchmarks show competitive performance and consistent improvements over the matched empirical risk minimization (ERM) baseline, although the size of the gain varies across datasets and target domains.