As image–text posts have become a dominant form of social media communication, understanding how the two modalities jointly convey meaning remains a key challenge in multimodal analysis. This study aims to examine whether image–text consistency is inherently multidimensional rather than reducible to a single similarity metric. Existing studies often reduce consistency to a single relevance score, which cannot capture semantic, emotional, and functional interactions. We construct a dataset of 28,650 multimodal posts and model image–text relationships along three dimensions: semantic consistency (CSC), emotional consistency (CEC), and informational matching consistency (IMC). Semantic and emotional alignment are measured using cross-modal representation and similarity computation, while IMC is defined through rule-based classification of informational roles. Results show that emotional consistency (CEC = 0.621) is higher than semantic consistency (CSC = 0.549, p<0.001), while 61.0% of posts maintain consistent informational orientation. These findings demonstrate that image–text consistency exhibits distinct cross-dimensional patterns that cannot be captured by single-metric approaches.
Xia et al. (Tue,) studied this question.