PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026SAGE Open6 citationsOpen Access

Translating Culture-Specific Items in The True Story of Ah Q: A Comparative Evaluation of GPT-4o, KIMI, and Google Translate Under Multimodal Prompts

View Full Paper
QWQiufen WangMAMansour AminiCYChen Yang

Key Points

  • The aim is to assess how well different translation models handle culture-specific items in a literary context.
  • Compared translation quality of GPT-4o, KIMI, DeepL, and Google Translate.
  • Evaluated performance using Cultural Adequacy, Linguistic Naturalness, and Terminology Accuracy.
  • Utilized a dataset of 30 culture-specific items from The True Story of Ah Q with human reference translations.
  • Implemented CLIP analysis for image-text alignment and qualitative assessment of translations.
  • GPT-4o with images achieved the highest scores in all evaluation dimensions.
  • Visual prompts enhanced translation quality significantly, particularly for GPT-4o.
  • KIMI showed some improvement but not statistically significant compared to human judgments.
  • DeepL and Google Translate performed worse than the multimodal systems and simplified cultural content more often.

Abstract

This study evaluates how multimodal large language models translate Chinese culture-specific items by comparing GPT-4o, KIMI, DeepL, and Google Translate. Building on a curated dataset of 30 CSIs from The True Story of Ah Q, each paired with two human reference translations and culturally relevant images, we assess three dimensions of quality: Cultural Adequacy, Linguistic Naturalness, and Terminology Accuracy. GPT-4o and KIMI are tested under text only and image plus text conditions, while DeepL and Google Translate serve as unimodal baselines. Methods combine expert ratings with CLIP image–text alignment and post hoc qualitative analysis of cultural simplification. Results show that visual prompts significantly improve human-rated quality, with GPT-4o (image plus text) achieving the highest scores across all dimensions. CLIP analysis indicates a significant gain for GPT-4o with images, while KIMI’s CLIP gain is not statistically significant, mirroring but not fully matching human judgments. DeepL and Google Translate trail the multimodal systems on all human-rated dimensions and exhibit more frequent cultural simplification. The study contributes a replicable multimodal evaluation framework and underscores the importance of visual context for culturally sensitive translation in literary settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/6980fd60c1c9540dea80f19ahttps://doi.org/10.1177/21582440251412225
Ask AI
Helpful
Bookmark
Share
View Full Paper