PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 19, 2026Biomimetics0 citationsOpen Access

HOIMamba: Bidirectional State-Space Modeling for Monocular 3D Human–Object Interaction Reconstruction

View Full Paper
JZJinsong ZhangYLYuqin Lin

Key Points

  • The aim is to enhance 3D human-object interaction reconstruction by addressing issues in traditional modeling approaches.
  • Developed HOIMamba as a bidirectional state-space modeling framework.
  • Introduced a multi-scale state-space module (MSSM) for capturing interaction dependencies.
  • Implemented a spatial-channel grouped SSM (SCSSM) for factorizing interaction models.
  • Evaluated the model on public datasets BEHAVE and InterCap using Chamfer distance.
  • Reduced human Chamfer distance by 8.6% on the BEHAVE dataset.
  • Improved contact recall by 13.5% compared to a strong Transformer-based baseline.
  • Achieved similar gains on the InterCap dataset.
  • Confirmed the efficacy of state-space modeling and bidirectional interaction reasoning through ablation studies.

Abstract

Monocular 3D human–object interaction (HOI) reconstruction requires jointly recovering articulated human geometry, object pose, and physically plausible contact from a single RGB image. While recent token-based methods commonly employ dense self-attention to capture global dependencies, isotropic all-to-all mixing tends to entangle spatial-geometric cues (e.g., contact locality) with channel-wise semantic cues (e.g., action/affordance), and provides limited control for representing directional and asymmetric physical influence between humans and objects. This paper presents HOIMamba, a state-space sequence modeling framework that reformulates HOI reconstruction as bidirectional, multi-scale interaction state inference. Instead of relying on symmetric correlation aggregation, HOIMamba uses structured state evolution to propagate interaction evidence. We introduce a multi-scale state-space module (MSSM) to capture interaction dependencies spanning local contact details and global body–object coordination. Building on MSSM, we propose a spatial-channel grouped SSM (SCSSM) block that factorizes interaction modeling into a spatial pathway for geometric/contact dependencies and a channel pathway for semantic/functional correlations, followed by gated fusion. HOIMamba further performs explicit bidirectional propagation between human and object states to better reflect asymmetric reciprocity in physical interactions. We evaluate HOIMamba on two public benchmarks, BEHAVE and InterCap, using Chamfer distance for human/object meshes and contact precision/recall induced by reconstructed geometry. HOIMamba achieves consistent improvements over representative prior methods. On the BEHAVE dataset, it reduces human Chamfer distance by 8.6% and improves contact recall by 13.5% compared to the strongest Transformer-based baseline, with similar gains observed on the InterCap dataset. Ablation studies on BEHAVE verify the contributions of state-space modeling, multi-scale inference, spatial-channel factorization, and bidirectional interaction reasoning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69bb9357496e729e62981668https://doi.org/10.3390/biomimetics11030214
Ask AI
Helpful
Bookmark
Share
View Full Paper