DOI: 10.3390/biomimetics11100687 ISSN: 2313-7673

SET-HOI: Skeleton-Enhanced Transformer with Parallel Fusion for HOI Detection

Rui Xiong, Bohong Wu, Qing Gao

Human–Object Interaction (HOI) detection seeks to identify the relationships between humans and objects in images that play a pivotal role in high-level vision tasks and biomimetic systems, such as biomimetic robotics, intelligent prosthetics, and embodied AI for intent understanding. However, existing approaches that rely solely on visual appearance features are susceptible to background interference, which adversely affects detection accuracy. Inspired by the biomimetic dual-stream mechanism of biological visual perception, we propose a Transformer-based HOI detection model with a dual-branch parallel fusion architecture, incorporating a skeleton topology branch alongside the standard image branch. The skeleton branch leverages a Graph Convolutional Network (GCN) to explicitly model spatial–kinematic relationships between human keypoints and objects, providing fine-grained topological priors to complement visual appearance features before joint decoding. Our approach achieves 51.6% Mean Average Precision (mAP) on the Verbs in Common Objects in Context (V-COCO) dataset, outperforming the Transformer-based HOI baseline by 2.7 mAP points, highlighting its effectiveness in enhancing perceptual robustness for biomimetic interactive applications.