PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 11, 20240 citationsOpen Access

An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set

View Full Paper
CAChaoyi Ai

Key Points

Key points are not available for this paper at this time.

Abstract

Human-Object Interaction (HOI) aims to identify the pairs of humans and objects in images and to recognize their relationships, ultimately forming human, object, verb triplets. Under default settings, HOI performance is nearly saturated, with many studies focusing on long-tail distribution and zero-shot/few-shot scenarios. Let us consider an intriguing problem: ``What if there is only test dataset without training dataset, using multimodal visual foundation model in a training-free manner? '' This study uses two experimental settings: grounding truth and random arbitrary combinations. We get some interesting conclusion and find that the open vocabulary capabilities of the multimodal visual foundation model are not yet fully realized. Additionally, replacing the feature extraction with grounding DINO further confirms these findings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chaoyi Ai (2024) studied this question.

synapsesocial.com/papers/68e5cb6fb6db643587562474https://doi.org/10.48550/arxiv.2408.05772
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Toward Open-Set Human Object Interaction Detection2024 · 7 citations
  2. 2Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection2024
  3. 3CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods2025
  4. 4Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration2024
  5. 5Contextual Human Object Interaction Understanding from Pre-Trained Large Language Model2024 · 7 citations