Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Karthika, N J, Suryanarayanan, Keerthana, Purohit, Jahanvi, Ramakrishnan, Ganesh, Singla, Jitin, Gourishetty, Anil Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
by: Maheshwari, Ayush, et al.
Published: (2023)
by: Maheshwari, Ayush, et al.
Published: (2023)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024)
by: J, Karthika N, et al.
Published: (2024)
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Debiasing Large Language Models toward Social Factors in Online Behavior Analytics through Prompt Knowledge Tuning
by: Salemi, Hossein, et al.
Published: (2026)
by: Salemi, Hossein, et al.
Published: (2026)
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
by: Karthika, N J, et al.
Published: (2025)
by: Karthika, N J, et al.
Published: (2025)
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024)
by: Maheshwari, Ayush, et al.
Published: (2024)
Semantically Cohesive Word Grouping in Indian Languages
by: Karthika, N J, et al.
Published: (2025)
by: Karthika, N J, et al.
Published: (2025)
Automatic Speech Recognition for Hindi
by: Saha, Anish, et al.
Published: (2024)
by: Saha, Anish, et al.
Published: (2024)
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
by: Kamath, Anusha, et al.
Published: (2025)
by: Kamath, Anusha, et al.
Published: (2025)
ACADATA: Parallel Dataset of Academic Data for Machine Translation
by: Lacunza, Iñaki, et al.
Published: (2025)
by: Lacunza, Iñaki, et al.
Published: (2025)
MorphTok: Morphologically Grounded Tokenization for Indian Languages
by: Brahma, Maharaj, et al.
Published: (2025)
by: Brahma, Maharaj, et al.
Published: (2025)
Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis
by: Shetty, Ahan Prasannakumar
Published: (2025)
by: Shetty, Ahan Prasannakumar
Published: (2025)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
YOLO-LAN: Precise Polyp Detection via Optimized Loss, Augmentations and Negatives
by: Gupta, Siddharth, et al.
Published: (2025)
by: Gupta, Siddharth, et al.
Published: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
by: Akkiraju, Bhavana, et al.
Published: (2025)
by: Akkiraju, Bhavana, et al.
Published: (2025)
Neural Compound-Word (Sandhi) Generation and Splitting in Sanskrit Language
by: Dave, Sushant, et al.
Published: (2020)
by: Dave, Sushant, et al.
Published: (2020)
A Benchmark Corpus and Neural Approach for Sanskrit Derivative Nouns Analysis
by: Singh, Arun Kumar, et al.
Published: (2020)
by: Singh, Arun Kumar, et al.
Published: (2020)
Using Deep Learning to Generate Semantically Correct Hindi Captions
by: Khan, Wasim Akram, et al.
Published: (2026)
by: Khan, Wasim Akram, et al.
Published: (2026)
Breaking Language Barriers: A Question Answering Dataset for Hindi and Marathi
by: Sabane, Maithili, et al.
Published: (2023)
by: Sabane, Maithili, et al.
Published: (2023)
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
by: Singh, Pooja, et al.
Published: (2025)
by: Singh, Pooja, et al.
Published: (2025)
Accent Placement Models for Rigvedic Sanskrit Text
by: P, Akhil Rajeev, et al.
Published: (2025)
by: P, Akhil Rajeev, et al.
Published: (2025)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
HindiLLM: Large Language Model for Hindi
by: Chouhan, Sanjay, et al.
Published: (2024)
by: Chouhan, Sanjay, et al.
Published: (2024)
SinFoS: A Parallel Dataset for Translating Sinhala Figures of Speech
by: Sofalas, Johan, et al.
Published: (2026)
by: Sofalas, Johan, et al.
Published: (2026)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
by: Nehrdich, Sebastian, et al.
Published: (2024)
by: Nehrdich, Sebastian, et al.
Published: (2024)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
by: Kumar, Sujeet, et al.
Published: (2025)
by: Kumar, Sujeet, et al.
Published: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
KazParC: Kazakh Parallel Corpus for Machine Translation
by: Yeshpanov, Rustem, et al.
Published: (2024)
by: Yeshpanov, Rustem, et al.
Published: (2024)
Automatic Speech Recognition for Sanskrit with Transfer Learning
by: Sadhukhan, Bidit, et al.
Published: (2025)
by: Sadhukhan, Bidit, et al.
Published: (2025)
XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
by: Liu, Jingxuan, et al.
Published: (2026)
by: Liu, Jingxuan, et al.
Published: (2026)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
EmoMix-3L: A Code-Mixed Dataset for Bangla-English-Hindi Emotion Detection
by: Raihan, Nishat, et al.
Published: (2024)
by: Raihan, Nishat, et al.
Published: (2024)
Sanskrit Knowledge-based Systems: Annotation and Computational Tools
by: Terdalkar, Hrishikesh
Published: (2024)
by: Terdalkar, Hrishikesh
Published: (2024)
Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
by: Choudhury, Monojit, et al.
Published: (2025)
by: Choudhury, Monojit, et al.
Published: (2025)
Similar Items
-
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
by: Maheshwari, Ayush, et al.
Published: (2023) -
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024) -
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026) -
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024) -
Debiasing Large Language Models toward Social Factors in Online Behavior Analytics through Prompt Knowledge Tuning
by: Salemi, Hossein, et al.
Published: (2026)