DOI: 10.3390/s26196083 ISSN: 1424-8220

Prompt-Controlled Multi-Modal Tuning for Person-Specific Few-Shot Micro-Expression Recognition

Ruiqi Wang, Kerong Li, Jiateng Liu, Tianxiang Cao, Yingtian Yu, Hengcan Shi, Yaonan Wang

Micro-expression recognition (MER) is a challenging step in various multimedia applications, such as media understanding and human–computer interaction, as it can reveal genuine human emotions. However, traditional MER often overlooks person-specific facial nuances, limiting generalization and personalized adaptability in practical applications. In this paper, we propose a novel person-specific few-shot MER benchmark that uses only an image/video of the target person as a reference to extend the boundaries of generalization and personalized adaptability in MER. A Prompt-Controlled Multi-Modal Tuning (PCMMT) framework is presented to tackle this benchmark. We introduce a prompt bottleneck mechanism in PCMMT that leverages text prompts and the CLIP multi-modal embedding space to bridge the query and reference visual inputs. Our prompt bottleneck also controls the flow of information across modalities to extract person-specific, subtle micro-expression motions. Moreover, adapter groups are designed at the layer level, with multi-modal tokens fine-tuned separately to refine person-specific cues and enhance generalization. The experimental results show that the proposed person-specific few-shot setting achieves better personalized generalization than traditional MER settings, and our PCMMT outperforms previous state-of-the-art models across various MER benchmarks.