PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 11, 20240 citationsOpen Access

Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration

View Full Paper
WXWeiying XueQLQi LiuQXQiwei Xiong

Key Points

Key points are not available for this paper at this time.

Abstract

Human-object interaction (HOI) detection aims to locate human-object pairs and identify their interaction categories in images. Most existing methods primarily focus on supervised learning, which relies on extensive manual HOI annotations. In this paper, we propose a novel framework, termed Knowledge Integration to HOI (KI2HOI), that effectively integrates the knowledge of visual-language model to improve zero-shot HOI detection. Specifically, the verb feature learning module is designed based on visual semantics, by employing the verb extraction decoder to convert corresponding verb queries into interaction-specific category representations. We develop an effective additive self-attention mechanism to generate more comprehensive visual representations. Moreover, the innovative interaction representation decoder effectively extracts informative regions by integrating spatial and visual feature information through a cross-attention mechanism. To deal with zero-shot learning in low-data, we leverage a priori knowledge from the CLIP text encoder to initialize the linear classifier for enhanced interaction understanding. Extensive experiments conducted on the mainstream HICO-DET and V-COCO datasets demonstrate that our model outperforms the previous methods in various zero-shot and full-supervised settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xue et al. (2024) studied this question.

synapsesocial.com/papers/68e74ba6b6db6435876c4893https://doi.org/10.48550/arxiv.2403.07246
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer2024 · 1 citations
  2. 2Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model2024 · 7 citations
  3. 3Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model2024
  4. 4Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection2025
  5. 5Locality-Aware Zero-Shot Human-Object Interaction Detection2025