PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 2026Pattern Recognition and Image Analysis3 citations

Analyzing Image Patterns and Generating Text: Advances in Multilingual Vision-Language Transformers

View Full Paper
STSergo TsiramuaHMHamlet MeladzeDBDavit Bitmalkishev

Key Points

  • This research aims to develop an artificial neural network that generates text and audio descriptions from images in multiple languages.
  • Developed the model 'Marta' utilizing a language model variant of text-to-text transfer transformer.
  • Integrated base-sized vision transformer (ViT-B-16) for image feature extraction using over 200,000 'image/description' pairs in Georgian.
  • Explored approaches to enhance multilingual capabilities, including encoder/decoder modifications and model retraining.
  • Identified advantages and weaknesses of various neural network models in processing multilingual tasks.
  • Highlighted the limitation of existing models focusing primarily on English language tasks.
  • Proposed technical solutions to support multilingual image and text processing effectively.

Abstract

Creating an artificial neural network capable of generating textual and audio descriptions of graphical images in multiple languages holds significant value from technological, social, economic, and humanitarian perspectives. This paper discusses neural networks and machine learning models focused on processing graphical images (pictures) and producing textual descriptions that can be used to solve multilingual tasks. These include convolutional neural networks, recurrent neural networks, its improved version called long short-term memory, as well as transformer-based models, bidirectional encoder representations from transformers, its multilingual version, and bootstrapped language-image pretraining. The paper presents a comparative analysis of these models, allowing the identification of their advantages and weaknesses when performing multilingual and visual tasks. Special attention is given to the bootstrapped language-image pretraining model, which is designed for simultaneous processing of text and images within a unified framework. Its main limitation is its focus on the English language, which poses a challenge for non-English language tasks. In the present paper is explored several approaches to overcoming this limitation: (1) integrating a translator at the output stage; (2) replacing the encoder/decoder with implementations that support the target language; (3) retraining the existing model using multilingual data. Each method comes with specific technical challenges, the analysis of which and corresponding recommendations are presented in the conclusion of the paper. To address this issue, we developed the model with name “Marta.” This model is based on the language model, which is a variant of the text-to-text transfer transformer model. To extract image features, the model uses base-sized vision transformer model that processes images as 16 × 16 patches (ViT-B-16), which is already trained for image feature recognition. Over 200 000 “image/description” pairs for training are already prepared in Georgian, so that the language model can learn how to handle the tensors output by the vision transformer model, that is, to correctly transform them into verbal interpretations and learn the logical process.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tsiramua et al. (2025) studied this question.

synapsesocial.com/papers/69f5939871405d493affe9f1https://doi.org/10.1134/s1054661825700750
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Bridging Languages through Images: A Multilingual Text-to-Image Synthesis Approach2024
  2. 2Applying Transformer-Based Neural Networks to Corpus Linguistic Data for Predictive Text Generation in Multilingual Environments2024
  3. 3Generative and Discriminative Models in Multimodal AI: An Analysis of Vision-Language Tasks2025
  4. 4Deep learning–driven image captioning: Progress through transformers and large language models2026 · 1 citations
  5. 5Investigation on task effect analysis and optimization strategy of multimodal large model based on Transformers architecture for various languages2024 · 1 citations