Monocular Interacting-Hand Reconstruction with Multimodal Context Fusion and Spatial Attention
Cheng Jin, Yiting Cao, Junfeng Zhu, Te Li, Hanzhang Wang, Yuchun Fang3D hand reconstruction from monocular RGB images has attracted increasing attention due to its low cost and ease of deployment. However, accurately reconstructing interacting hands remains challenging, primarily because mutual occlusion leads to missing visual evidence, while image cropping weakens global spatial consistency. To address these issues, we propose the Context-Aware Interacting Hand Reconstruction Network (CANet), a monocular interacting-hand reconstruction framework that integrates multimodal context fusion with structured spatial attention. Specifically, CANet leverages estimated depth maps, edge contours, and center heatmaps to retain global contextual cues and guide reconstruction in occluded regions. A hierarchical spatial attention module further enhances spatial consistency by separately modeling intra-hand structural dependencies and inter-hand interactions, enabling more coherent reasoning about complex hand poses. Experiments across multiple public benchmarks demonstrate that CANet consistently improves the accuracy of both hand pose estimation and mesh reconstruction.