PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

Behind Maya: Building a Multilingual Vision Language Model

View Full Paper
NAN. M. AlamNorth Carolina State UniversityKKKarthik Reddy KanjulaCoherent (United States)SGSurya GuthikondaCoherent (United States)

Key Points

  • Maya significantly improves performance for low-resource languages, expanding access to vision-language tasks for diverse cultures.
  • The model supports eight languages, utilizing a dataset that enhances cultural and linguistic understanding, addressing existing gaps.
  • Maya is built on the LLaVA pretraining dataset, incorporating a multilingual framework that enables better language support in vision tasks.
  • This initiative highlights the necessity of inclusivity in machine learning frameworks, ensuring equitable performance across linguistic groups.

Abstract

In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken languages but lack performance on low-resource languages and varied cultural contexts. To address these limitations, we introduce Maya, an open-source Multilingual VLM. Our contributions are: 1) a multilingual image-text pretraining dataset in eight languages, based on the LLaVA pretraining dataset; and 2) a multilingual image-text model supporting these languages, enhancing cultural and linguistic comprehension in vision-language tasks. Code available at https://github.com/nahidalam/maya.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alam et al. (2025) studied this question.

synapsesocial.com/papers/68f147cc724575985c3fd105https://doi.org/10.48550/arxiv.2505.08910
Ask AI
Helpful
Bookmark
Share
View Full Paper