PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 29, 20260 citationsOpen Access

Embedding Model Selection for Domain-Specific Retrieval-Augmented Generation: A Comparative Study on Indian Cultural Heritage Corpora

View Full Paper
PAPriyanka AsthanaLMLaxmi Shankar MauryaHKHarshita Kushwaha

Key Points

  • This research aims to assess the effectiveness of different sentence transformer models for retrieval-augmented generation in the context of Indian cultural heritage.
  • Evaluated three sentence transformer architectures: all-MiniLM-L6-v2, all-mpnet-base-v2, and paraphrase-multilingual-MiniLM-L12-v2.
  • Conducted experiments on a corpus of 15 documents containing 2,876 chunks, focusing on English-language retrieval.
  • Implemented processes including document ingestion, chunking, FAISS vector indexing, and grounded generation with LLaMA 3.1 8B.
  • Monolingual sentence transformer models outperformed multilingual alternatives for English-language retrieval.
  • The ranking of embedding models was found to reverse with varying corpus sizes, indicating the need for evaluation at large scales.
  • Manually validated question-answer evaluation demonstrated differences in performance across transformer architectures.

Abstract

This paper presents Kashivani, a domain-specific Retrieval-Augmented Generation (RAG) framework for Indian cultural heritage knowledge retrieval focused on Varanasi, India. The study evaluates three sentence transformer architectures, namely: all-MiniLM-L6-v2, all-mpnet-base-v2, and paraphrase-multilingual-MiniLM-L12-v2 across different corpus scales using manually validated question-answer evaluation. Experiments conducted on a corpus of 15 cultural heritage documents containing 2,876 chunks demonstrate that monolingual sentence transformer models consistently outperform multilingual alternatives for English-language Indian cultural heritage retrieval. The study further shows that embedding model rankings can reverse under corpus scaling conditions, highlighting the importance of evaluating retrieval systems at deployment-scale corpus sizes. The complete implementation includes document ingestion, chunking, FAISS vector indexing, semantic retrieval, and grounded generation using LLaMA 3.1 8B.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Asthana et al. (2026) studied this question.

synapsesocial.com/papers/6a192f2dfab5b468c44189c1https://doi.org/10.5281/zenodo.20399129
Ask AI
Helpful
Bookmark
Share
View Full Paper