PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 2025IEEE Transactions on Image Processing1 citations

Hierarchical Multimodal Knowledge Matching for Training-Free Open-Vocabulary Object Detection

View Full Paper
QMQisen MaYHYan HuangZLZikun Liu

Key Points

  • The hierarchical multimodal knowledge matching method improves detection results for novel categories.
  • Using limited category-specific images, it builds object and attribute prototype knowledge for better representation.
  • The proposed training-free approach allows seamless integration into existing open-vocabulary object detection models.
  • Extensive experiments verify significant performance improvements across various datasets and model architectures.

Abstract

Open-Vocabulary Object Detection (OVOD) aims to leverage the generalization capabilities of pre-trained vision-language models for detecting objects beyond the trained categories. Existing methods mostly focus on supervised learning strategies based on available training data, which might be suboptimal for data-limited novel categories. To tackle this challenge, this paper presents a Hierarchical Multimodal Knowledge Matching method (HMKM) to better represent novel categories and match them with region features. Specifically, HMKM includes a set of object prototype knowledge that is obtained using limited category-specific images, acting as off-the-shelf category representations. In addition, HMKM also includes a set of attribute prototype knowledge to represent key attributes of categories at a fine-grained level, with the goal to distinguish one category from its visually similar ones. During inference, two sets of object and attribute prototype knowledge are adaptively combined to match categories with region features. The proposed HMKM is training-free and can be easily integrated as a plug-and-play module into existing OVOD models. Extensive experiments demonstrate that our HMKM significantly improves the performance when detecting novel categories across various backbones and datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ma et al. (2025) studied this question.

synapsesocial.com/papers/68f10ecee6a12fd0428998d2https://doi.org/10.1109/tip.2025.3618408
Ask AI
Helpful
Bookmark
Share
View Full Paper