PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025Information, computing and intelligent systems3 citationsOpen Access

A Multimodal Retrieval-Augmented Generation System with ReAct Agent Logic for Multi-Hop Reasoning

View Full Paper
DYDenys YuvzhenkoVCViacheslaw ChymshyrVSVolodymyr Shymkovych

Key Points

  • The proposed system achieved an answer accuracy of 92%, indicating its effectiveness in generating relevant responses.
  • Using a specialized test set, the system managed to cover all data types completely, ensuring comprehensive multimodal integration.
  • Experimental methods included semantic vector search and multi-hop planning, validating the ReAct logic's efficiency and response speed.
  • The study found that the multimodal agent-based approach provides significant advantages over traditional text-only systems.

Abstract

The rapid advancement of generative artificial intelligence models significantly influences modern methods of information processing and user interactions with information systems. One of the promising areas in this domain is Retrieval-Augmented Generation (RAG), which combines generative models with information retrieval methods to enhance the accuracy and relevance of responses. However, most existing RAG systems primarily focus on textual data, which does not meet contemporary needs for multimodal information processing (text, images, tables). The research object of this work is a multimodal RAG system based on ReAct agent logic, capable of multi-hop reasoning. The main emphasis is placed on integrating textual, graphical, and tabular information to generate accurate, complete, and relevant responses. The system's implementation utilized the ChromaDB vector storage, the OpenAI embedding generation model (text-embedding-ada-002), and the GPT-4 language model. The purpose of the study is the development, deployment, and empirical evaluation of the proposed multimodal RAG system based on the ReAct agent approach, capable of effectively integrating diverse knowledge sources into a unified informational context. The experimental evaluation utilized the Global Tuberculosis Report 2024 by the World Health Organization, containing various textual, graphical, and tabular data. A specialized test set of 50 queries (30 textual, 10 tabular, 10 graphical) was created for empirical analysis, allowing comprehensive testing of all aspects of multimodal integration. The research employed methods such as semantic vector search, multi-hop agent-based planning with ReAct logic, and evaluations of answer accuracy, answer recall, and response latency. Additionally, an analysis of response speed dependence on query volume was conducted. The obtained results confirmed the high efficiency of the proposed approach. The system demonstrated an answer accuracy of 92%, answer recall of 89%, and ensured complete (100%) coverage of all data types. The average response time was approximately 5 seconds, meeting interactive system requirements. Optimal parameters were experimentally determined (for example, parameter k = 6, classification threshold 0.35, and up to three reasoning iterations), ensuring the best balance among completeness, speed, and operational efficiency. The study's findings highlighted significant advantages of the multimodal agent-based approach compared to traditional textual RAG solutions, confirming the promising direction for further research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yuvzhenko et al. (2025) studied this question.

synapsesocial.com/papers/68d46fc631b076d99fa69bcfhttps://doi.org/10.20535/2786-8729.6.2025.330777
Ask AI
Helpful
Bookmark
Share
View Full Paper