Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karthika, N J, Suryanarayanan, Keerthana, Purohit, Jahanvi, Ramakrishnan, Ganesh, Singla, Jitin, Gourishetty, Anil Kumar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914421095268352
author Karthika, N J
Suryanarayanan, Keerthana
Purohit, Jahanvi
Ramakrishnan, Ganesh
Singla, Jitin
Gourishetty, Anil Kumar
author_facet Karthika, N J
Suryanarayanan, Keerthana
Purohit, Jahanvi
Ramakrishnan, Ganesh
Singla, Jitin
Gourishetty, Anil Kumar
contents We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this corpus aggregates data from diverse sources covering contemporary materials, including spoken tutorials, children's magazines, radio conversations, and instruction materials. We benchmark this new dataset by fine-tuning three complementary models - ByT5, NLLB and IndicTrans-v2, to demonstrate its utility. Our experiments demonstrate that models trained on the Samasamayik corpus achieve significant performance gains on in-domain test data, while achieving comparable performance on other widely used test sets, establishing a strong new performance baseline for contemporary Hindi-Sanskrit translation. Furthermore, a comparative analysis against existing corpora reveals minimal semantic and lexical overlap, confirming the novelty and non-redundancy of our dataset as a robust new resource for low-resource Indic language MT.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24307
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
Karthika, N J
Suryanarayanan, Keerthana
Purohit, Jahanvi
Ramakrishnan, Ganesh
Singla, Jitin
Gourishetty, Anil Kumar
Computation and Language
We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this corpus aggregates data from diverse sources covering contemporary materials, including spoken tutorials, children's magazines, radio conversations, and instruction materials. We benchmark this new dataset by fine-tuning three complementary models - ByT5, NLLB and IndicTrans-v2, to demonstrate its utility. Our experiments demonstrate that models trained on the Samasamayik corpus achieve significant performance gains on in-domain test data, while achieving comparable performance on other widely used test sets, establishing a strong new performance baseline for contemporary Hindi-Sanskrit translation. Furthermore, a comparative analysis against existing corpora reveals minimal semantic and lexical overlap, confirming the novelty and non-redundancy of our dataset as a robust new resource for low-resource Indic language MT.
title Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2603.24307