Multi‐Modal Perception in Dynamic Occlusion Scenarios: A Human Pose Estimation Approach in Human‐Robot Collaboration
Lijie Zhou, Hongyu Wang, Xingqi Li, Bingchen Song, Zehai Huang, Jia ZhangABSTRACT
The human pose estimation technology based on robotic vision sensing systems has emerged as a critical perception approach for collision avoidance and safety assurance in human‐robot collaboration (HRC). In dynamic occlusion scenarios, human pose estimation primarily relies on rapidly evolving deep learning models, yet it still faces challenges such as degraded recognition accuracy under prolonged occlusion and the high time cost of data collection and annotation. To address this issue, we proposed a novel skeletal pose and minimum‐distance fusion (SPMF) approach which aims to achieve robust human pose estimation in dynamically occluded human‐robot collaboration (HRC) scenarios. Firstly, the 3D coordinates of the joints were pre‐estimated using the OpenPose‐based skeletal model, in which the human body was represented by capsules at the joint positions. Following, the minimum human‐robot distance was calculated via the Gilbert‐Johnson‐Keerthi (GJK) algorithm by combining the capsule model of the robot. In addition, a framework for human postural discrimination was established by combining the joint prediction confidence and minimum distance parameters, which can judge the reliability of pre‐estimated pose data for occluded points. Finally, an algorithm for correcting misjudged joints was proposed which reconstructed the depth values for fully occluded joints. The experiments conducted on a refined Human3.6M‐Occluded‐SV database show that the proposed method substantially outperforms other fusion methods under severe occlusion. Furthermore, the experiments on real HRC scenarios demonstrate that the pose estimation framework can achieve real‐time collision detection under robot occlusion and deliver accurate pose recognition for industrial safety applications.