Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mekonnen, Kidist Amde, Alemneh, Yosef Worku, de Rijke, Maarten
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916786678530048
author Mekonnen, Kidist Amde
Alemneh, Yosef Worku
de Rijke, Maarten
author_facet Mekonnen, Kidist Amde
Alemneh, Yosef Worku
de Rijke, Maarten
contents Neural retrieval methods using transformer-based pre-trained language models have advanced multilingual and cross-lingual retrieval. However, their effectiveness for low-resource, morphologically rich languages such as Amharic remains underexplored due to data scarcity and suboptimal tokenization. We address this gap by introducing Amharic-specific dense retrieval models based on pre-trained Amharic BERT and RoBERTa backbones. Our proposed RoBERTa-Base-Amharic-Embed model (110M parameters) achieves a 17.6% relative improvement in MRR@10 and a 9.86% gain in Recall@10 over the strongest multilingual baseline, Arctic Embed 2.0 (568M parameters). More compact variants, such as RoBERTa-Medium-Amharic-Embed (42M), remain competitive while being over 13x smaller. Additionally, we train a ColBERT-based late interaction retrieval model that achieves the highest MRR@10 score (0.843) among all evaluated models. We benchmark our proposed models against both sparse and dense retrieval baselines to systematically assess retrieval effectiveness in Amharic. Our analysis highlights key challenges in low-resource settings and underscores the importance of language-specific adaptation. To foster future research in low-resource IR, we publicly release our dataset, codebase, and trained models at https://github.com/kidist-amde/amharic-ir-benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval
Mekonnen, Kidist Amde
Alemneh, Yosef Worku
de Rijke, Maarten
Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
68T50 (Primary), 68T05 (Secondary)
H.3.3; H.3.1; I.2.7
Neural retrieval methods using transformer-based pre-trained language models have advanced multilingual and cross-lingual retrieval. However, their effectiveness for low-resource, morphologically rich languages such as Amharic remains underexplored due to data scarcity and suboptimal tokenization. We address this gap by introducing Amharic-specific dense retrieval models based on pre-trained Amharic BERT and RoBERTa backbones. Our proposed RoBERTa-Base-Amharic-Embed model (110M parameters) achieves a 17.6% relative improvement in MRR@10 and a 9.86% gain in Recall@10 over the strongest multilingual baseline, Arctic Embed 2.0 (568M parameters). More compact variants, such as RoBERTa-Medium-Amharic-Embed (42M), remain competitive while being over 13x smaller. Additionally, we train a ColBERT-based late interaction retrieval model that achieves the highest MRR@10 score (0.843) among all evaluated models. We benchmark our proposed models against both sparse and dense retrieval baselines to systematically assess retrieval effectiveness in Amharic. Our analysis highlights key challenges in low-resource settings and underscores the importance of language-specific adaptation. To foster future research in low-resource IR, we publicly release our dataset, codebase, and trained models at https://github.com/kidist-amde/amharic-ir-benchmarks.
title Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval
topic Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
68T50 (Primary), 68T05 (Secondary)
H.3.3; H.3.1; I.2.7
url https://arxiv.org/abs/2505.19356