Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Maheshwari, Ayush, Gupta, Ashim, Krishna, Amrith, Singh, Atul Kumar, Ramakrishnan, Ganesh, Kumar, G. Anil, Singla, Jitin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
by: Karthika, N J, et al.
Published: (2026)
by: Karthika, N J, et al.
Published: (2026)
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
by: M., Yashwanth, et al.
Published: (2025)
by: M., Yashwanth, et al.
Published: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026)
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024)
by: Maheshwari, Ayush, et al.
Published: (2024)
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
Mitrasamgraha: A Comprehensive Classical Sanskrit Machine Translation Dataset
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages
by: J, Karthika N, et al.
Published: (2024)
by: J, Karthika N, et al.
Published: (2024)
A Benchmark Corpus and Neural Approach for Sanskrit Derivative Nouns Analysis
by: Singh, Arun Kumar, et al.
Published: (2020)
by: Singh, Arun Kumar, et al.
Published: (2020)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
YOLO-LAN: Precise Polyp Detection via Optimized Loss, Augmentations and Negatives
by: Gupta, Siddharth, et al.
Published: (2025)
by: Gupta, Siddharth, et al.
Published: (2025)
Neural Compound-Word (Sandhi) Generation and Splitting in Sanskrit Language
by: Dave, Sushant, et al.
Published: (2020)
by: Dave, Sushant, et al.
Published: (2020)
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
by: Kumar, Sujeet, et al.
Published: (2025)
by: Kumar, Sujeet, et al.
Published: (2025)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
Chandomitra: Towards Generating Structured Sanskrit Poetry from Natural Language Inputs
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation
by: Akkiraju, Bhavana, et al.
Published: (2025)
by: Akkiraju, Bhavana, et al.
Published: (2025)
AdiBhashaa: A Community-Curated Benchmark for Machine Translation into Indian Tribal Languages
by: Singh, Pooja, et al.
Published: (2025)
by: Singh, Pooja, et al.
Published: (2025)
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments
by: Panda, Debasmita, et al.
Published: (2026)
by: Panda, Debasmita, et al.
Published: (2026)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System
by: Mayilvaghanan, Kawin, et al.
Published: (2026)
by: Mayilvaghanan, Kawin, et al.
Published: (2026)
CSSL: Contrastive Self-Supervised Learning for Dependency Parsing on Relatively Free Word Ordered and Morphologically Rich Low Resource Languages
by: Ray, Pretam, et al.
Published: (2024)
by: Ray, Pretam, et al.
Published: (2024)
Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries
by: Mayilvaghanan, Kawin, et al.
Published: (2025)
by: Mayilvaghanan, Kawin, et al.
Published: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
Anomaly Detection in Human Language via Meta-Learning: A Few-Shot Approach
by: Singla, Saurav, et al.
Published: (2025)
by: Singla, Saurav, et al.
Published: (2025)
MorphTok: Morphologically Grounded Tokenization for Indian Languages
by: Brahma, Maharaj, et al.
Published: (2025)
by: Brahma, Maharaj, et al.
Published: (2025)
Accent Placement Models for Rigvedic Sanskrit Text
by: P, Akhil Rajeev, et al.
Published: (2025)
by: P, Akhil Rajeev, et al.
Published: (2025)
A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
by: Kunz, Jenny, et al.
Published: (2026)
by: Kunz, Jenny, et al.
Published: (2026)
EEG-to-Text Translation: A Model for Deciphering Human Brain Activity
by: Murad, Saydul Akbar, et al.
Published: (2025)
by: Murad, Saydul Akbar, et al.
Published: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
by: Kamath, Anusha, et al.
Published: (2025)
by: Kamath, Anusha, et al.
Published: (2025)
Verifiable Natural Language to Linear Temporal Logic Translation: A Benchmark Dataset and Evaluation Suite
by: English, William H, et al.
Published: (2025)
by: English, William H, et al.
Published: (2025)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
by: Nehrdich, Sebastian, et al.
Published: (2024)
by: Nehrdich, Sebastian, et al.
Published: (2024)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
A Code Comprehension Benchmark for Large Language Models for Code
by: Havare, Jayant, et al.
Published: (2025)
by: Havare, Jayant, et al.
Published: (2025)
FAIR: Filtering of Automatically Induced Rules
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Japanese-English Sentence Translation Exercises Dataset for Automatic Grading
by: Miura, Naoki, et al.
Published: (2024)
by: Miura, Naoki, et al.
Published: (2024)
Similar Items
-
Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation
by: Karthika, N J, et al.
Published: (2026) -
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
by: M., Yashwanth, et al.
Published: (2025) -
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2026) -
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024) -
A Three-Pronged Approach to Cross-Lingual Adaptation with Multilingual LLMs
by: Singh, Vaibhav, et al.
Published: (2024)