3D hand reconstruction from monocular RGB images has attracted increasing attention due to its low cost and ease of deployment. However, accurately reconstructing interacting hands remains challenging, primarily because mutual occlusion leads to missing visual evidence, while image cropping weakens global spatial consistency. To address these issues, we propose the Context-Aware Interacting Hand Reconstruction Network (CANet), a monocular interacting-hand reconstruction framework that integrates multimodal context fusion with structured spatial attention. Specifically, CANet leverages estimated depth maps, edge contours, and center heatmaps to retain global contextual cues and guide reconstruction in occluded regions. A hierarchical spatial attention module further enhances spatial consistency by separately modeling intra-hand structural dependencies and inter-hand interactions, enabling more coherent reasoning about complex hand poses. Experiments across multiple public benchmarks demonstrate that CANet consistently improves the accuracy of both hand pose estimation and mesh reconstruction.
No takes yet. Share an insight, caveat, or question.
Jin et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: