Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nehrdich, Sebastian, Allport, David, Sellmer, Sven, Sandhan, Jivnesh, Jagadeeshan, Manoj Balaji, Goyal, Pawan, Kumar, Sujeet, Keutzer, Kurt
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915723056513024
author Nehrdich, Sebastian
Allport, David
Sellmer, Sven
Sandhan, Jivnesh
Jagadeeshan, Manoj Balaji
Goyal, Pawan
Kumar, Sujeet
Keutzer, Kurt
author_facet Nehrdich, Sebastian
Allport, David
Sellmer, Sven
Sandhan, Jivnesh
Jagadeeshan, Manoj Balaji
Goyal, Pawan
Kumar, Sujeet
Keutzer, Kurt
contents While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challenges such as poetic language, philosophical concepts, multi-layered metaphorical expressions, and more. Sanskrit literature is a prime example of this, as it combines a large number of such challenges in addition to inherent linguistic features like sandhi, compounding, and heavy morphology, which further complicate NLP downstream tasks. It spans multiple millennia of text production time as well as a large breadth of different domains, ranging from ritual formulas via epic narratives, philosophical treatises, poetic verses up to scientific material. As of now, there is a strong lack of publicly available resources that cover these different domains and temporal layers of Sanskrit. We therefore introduce Mitrasamgraha, a high-quality Sanskrit-to-English machine translation dataset consisting of 391,548 bitext pairs, more than four times larger than the largest previously available Sanskrit dataset Itih=asa. It covers a time period of more than three millennia and a broad range of historical Sanskrit domains. In contrast to web-crawled datasets, the temporal and domain annotation of this dataset enables fine-grained study of domain and time period effects on MT performance. We also release a validation set consisting of 5,587 and a test set consisting of 5,552 post-corrected bitext pairs. We conduct experiments benchmarking commercial and open models on this dataset and fine-tune NLLB and Gemma models on the dataset, showing significant improvements, while still recognizing significant challenges in the translation of complex compounds, philosophical concepts, and multi-layered metaphors. We also analyze how in-context learning on this dataset impacts the performance of commercial models
format Preprint
id arxiv_https___arxiv_org_abs_2601_07314
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
Nehrdich, Sebastian
Allport, David
Sellmer, Sven
Sandhan, Jivnesh
Jagadeeshan, Manoj Balaji
Goyal, Pawan
Kumar, Sujeet
Keutzer, Kurt
Computation and Language
While machine translation is regarded as a "solved problem" for many high-resource languages, close analysis quickly reveals that this is not the case for content that shows challenges such as poetic language, philosophical concepts, multi-layered metaphorical expressions, and more. Sanskrit literature is a prime example of this, as it combines a large number of such challenges in addition to inherent linguistic features like sandhi, compounding, and heavy morphology, which further complicate NLP downstream tasks. It spans multiple millennia of text production time as well as a large breadth of different domains, ranging from ritual formulas via epic narratives, philosophical treatises, poetic verses up to scientific material. As of now, there is a strong lack of publicly available resources that cover these different domains and temporal layers of Sanskrit. We therefore introduce Mitrasamgraha, a high-quality Sanskrit-to-English machine translation dataset consisting of 391,548 bitext pairs, more than four times larger than the largest previously available Sanskrit dataset Itih=asa. It covers a time period of more than three millennia and a broad range of historical Sanskrit domains. In contrast to web-crawled datasets, the temporal and domain annotation of this dataset enables fine-grained study of domain and time period effects on MT performance. We also release a validation set consisting of 5,587 and a test set consisting of 5,552 post-corrected bitext pairs. We conduct experiments benchmarking commercial and open models on this dataset and fine-tune NLLB and Gemma models on the dataset, showing significant improvements, while still recognizing significant challenges in the translation of complex compounds, philosophical concepts, and multi-layered metaphors. We also analyze how in-context learning on this dataset impacts the performance of commercial models
title Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
topic Computation and Language
url https://arxiv.org/abs/2601.07314