PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

View Full Paper
WZWenchao ZhangJTJiahe TianRHReuben He

Key Points

  • Generated images from text prompts often lack alignment with real-world knowledge, leading to misunderstandings.
  • Using the ABP benchmark, 8 popular T2I models were evaluated, revealing state-of-the-art models have significant alignment limitations.
  • The novel ABPScore metric correlates strongly with human judgment, providing a more comprehensive evaluation framework.
  • Inference-Time Knowledge Injection significantly enhanced model performance, improving ABPScore by approximately 43%.

Abstract

Recent text-to-image (T2I) generation models have advanced significantly, enabling the creation of high-fidelity images from textual prompts. However, existing evaluation benchmarks primarily focus on the explicit alignment between generated images and prompts, neglecting the alignment with real-world knowledge beyond prompts. To address this gap, we introduce Align Beyond Prompts (ABP), a comprehensive benchmark designed to measure the alignment of generated images with real-world knowledge that extends beyond the explicit user prompts. ABP comprises over 2,000 meticulously crafted prompts, covering real-world knowledge across six distinct scenarios. We further introduce ABPScore, a metric that utilizes existing Multimodal Large Language Models (MLLMs) to assess the alignment between generated images and world knowledge beyond prompts, which demonstrates strong correlations with human judgments. Through a comprehensive evaluation of 8 popular T2I models using ABP, we find that even state-of-the-art models, such as GPT-4o, face limitations in integrating simple real-world knowledge into generated images. To mitigate this issue, we introduce a training-free strategy within ABP, named Inference-Time Knowledge Injection (ITKI). By applying this strategy to optimize 200 challenging samples, we achieved an improvement of approximately 43% in ABPScore. The dataset and code are available in https://github.com/smile365317/ABP.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68da58d8c1728099cfd111b7https://doi.org/10.48550/arxiv.2505.18730
Ask AI
Helpful
Bookmark
Share
View Full Paper