Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 29, 2025Open Access

Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

View Full Paper
Ask AI
Bookmark
Share

Authors

KMKidist Amde MekonnenYAYosef Worku AlemnehMRMaarten de Rijke

Discussion

Loading...

Member takes

Overview

Proposed Amharic-specific models improve retrieval effectiveness, highlighting challenges in low-resource settings.

Key Points

  • RoBERTa-Base-Amharic-Embed outperforms multilingual baselines, achieving a 17.6% improvement in MRR@10.
  • The introduced models demonstrate significant gains in recall, with RoBERTa-Medium-Amharic-Embed being over 13x smaller yet competitive.
  • A ColBERT-based model achieves the highest MRR@10 score among evaluated models, showcasing advancements in dense retrieval.
  • Benchmarking against various retrieval baselines underlines the necessity for language-specific adaptation in low-resource conditions.

Cite This Study

Mekonnen et al. (2025) studied this question.

synapsesocial.com/papers/68da5a3ec1728099cfd1196fhttps://doi.org/10.48550/arxiv.2505.19356
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval2026 · 2 citations
  2. 2Transformer-based coreference resolution modeling for Amharic text2026
  3. 3News Classification in Low‐Resource Languages: Insights From Transformer and Baseline Models2026 · 1 citations
  4. 4Bidirectional Transformer-Based Neural Machine Translation for Amharic and Tigrinya: Bridging Morphological Complexity and Data Scarcity2025
  5. 5ColBERT-XM: A Modular Multi-Vector Representation Model for Zero-Shot Multilingual Information Retrieval2024 · 1 citations