Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 3, 2025Open Access

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters

View Full Paper
Ask AI
Bookmark
Share

Authors

ZCZijian ChenShandong UniversityWHWenjie HuaWuhan UniversityJLJinhao LiEast China Normal University

Discussion

Loading...

Member takes

Overview

Dataset PictOBI-20k evaluates visual decipherment in oracle bone characters, highlighting LMMs’ limitations with visual data.

Key Points

  • General large multimodal models show preliminary skills in visually deciphering oracle bone characters, yet they often fail to utilize visual information effectively.
  • The PictOBI-20k dataset features 20k images and over 15k multiple-choice questions, designed to enhance visual reasoning in LMMs.
  • Subjective annotations are used to clarify the consistency of reference points between human understanding and large multimodal models.
  • Findings suggest that language priors significantly limit LMMs in visual reasoning for understanding pictographic characters.

Cite This Study

Chen et al. (2025) studied this question.

synapsesocial.com/papers/68e02f3cf0e39f13e7fa255ehttps://doi.org/10.48550/arxiv.2509.05773
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1OBIMD: A Multi-modal Dataset for Contextual Interpretation of Oracle Bone Inscriptions2026 · 1 citations
  2. 2Detecting oracle bone inscriptions via pseudo-category labels2024 · 22 citations
  3. 3Leveraging progressive domain adaptation for unsupervised cross-domain oracle bone inscription recognition2026
  4. 4Prism-OBI: a novel framework for oracle bone inscription recognition via visual perception and feature decoupling2026 · 1 citations
  5. 5A text image dual conditional stable diffusion model for oracle bone inscription decipherment2025 · 5 citations