INDIC DIALECT: A Multi Task Benchmark to Evaluate and Translate in Indian Language Dialects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sharma, Tarun, Ravikiran, Manikandan, Behera, Sourava Kumar, Bhattacharya, Pramit, Bhattacharya, Arnab, Saluja, Rohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2024)
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2024)
Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025)
BanglaByT5: Byte-Level Modelling for Bangla
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025)
Vacaspati: A Diverse Corpus of Bangla Literature
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2023)
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2023)
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
PARAMANU-GANITA: Can Small Math Language Models Rival with Large Language Models on Mathematical Reasoning?
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)
Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
BhashaSetu: Cross-Lingual Knowledge Transfer from High-Resource to Extreme Low-Resource Languages
von: Maji, Subhadip, et al.
Veröffentlicht: (2026)
von: Maji, Subhadip, et al.
Veröffentlicht: (2026)
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
von: Karthika, N J, et al.
Veröffentlicht: (2025)
von: Karthika, N J, et al.
Veröffentlicht: (2025)
Semantically Cohesive Word Grouping in Indian Languages
von: Karthika, N J, et al.
Veröffentlicht: (2025)
von: Karthika, N J, et al.
Veröffentlicht: (2025)
Legal Judgment Reimagined: PredEx and the Rise of Intelligent AI Interpretation in Indian Courts
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
IBPS: Indian Bail Prediction System
von: Srivastava, Puspesh Kumar, et al.
Veröffentlicht: (2025)
von: Srivastava, Puspesh Kumar, et al.
Veröffentlicht: (2025)
CODET: A Benchmark for Contrastive Dialectal Evaluation of Machine Translation
von: Alam, Md Mahfuz Ibn, et al.
Veröffentlicht: (2023)
von: Alam, Md Mahfuz Ibn, et al.
Veröffentlicht: (2023)
LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
A Likelihood Ratio Test of Genetic Relationship among Languages
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2024)
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2024)
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2024)
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2024)
MorphTok: Morphologically Grounded Tokenization for Indian Languages
von: Brahma, Maharaj, et al.
Veröffentlicht: (2025)
von: Brahma, Maharaj, et al.
Veröffentlicht: (2025)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2025)
von: Akavarapu, V. S. D. S. Mahesh, et al.
Veröffentlicht: (2025)
Building pre-train LLM Dataset for the INDIC Languages: a case study on Hindi
von: Parida, Shantipriya, et al.
Veröffentlicht: (2024)
von: Parida, Shantipriya, et al.
Veröffentlicht: (2024)
AdiBhashaa: A Community-Curated Benchmark for Machine Translation into Indian Tribal Languages
von: Singh, Pooja, et al.
Veröffentlicht: (2025)
von: Singh, Pooja, et al.
Veröffentlicht: (2025)
Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
NyayaAnumana & INLegalLlama: The Largest Indian Legal Judgment Prediction Dataset and Specialized Language Model for Enhanced Decision Analysis
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2024)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
Towards Large Language Model driven Reference-less Translation Evaluation for English and Indian Languages
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
von: Bhattacharya, Amartya
Veröffentlicht: (2026)
von: Bhattacharya, Amartya
Veröffentlicht: (2026)
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
von: Blaschke, Verena, et al.
Veröffentlicht: (2025)
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
von: Mujadia, Vandan, et al.
Veröffentlicht: (2024)
TathyaNyaya and FactLegalLlama: Advancing Factual Judgment Prediction and Explanation in the Indian Legal Context
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
Crosslingual Optimized Metric for Translation Assessment of Indian Languages
von: Ahsan, Arafat, et al.
Veröffentlicht: (2025)
von: Ahsan, Arafat, et al.
Veröffentlicht: (2025)
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2023)
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2023)
MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages
von: Mahapatra, Sayan, et al.
Veröffentlicht: (2023)
von: Mahapatra, Sayan, et al.
Veröffentlicht: (2023)
A Gold Standard Dataset and Evaluation Framework for Depression Detection and Explanation in Social Media using LLMs
von: Bolegave, Prajval, et al.
Veröffentlicht: (2025)
von: Bolegave, Prajval, et al.
Veröffentlicht: (2025)
Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study
von: Krückl, Xaver Maria, et al.
Veröffentlicht: (2025)
von: Krückl, Xaver Maria, et al.
Veröffentlicht: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
von: Altakrori, Malik H., et al.
Veröffentlicht: (2025)
von: Altakrori, Malik H., et al.
Veröffentlicht: (2025)
Optimizing delivery for quick commerce factoring qualitative assessment of generated routes
von: Bhattacharya, Milon, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Milon, et al.
Veröffentlicht: (2025)
Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Debarpan, et al.
Veröffentlicht: (2025)
Segment First, Retrieve Better: Realistic Legal Search via Rhetorical Role-Based Queries
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging LLMs for Bangla Grammar Error Correction:Error Categorization, Synthetic Data, and Model Evaluation
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2024) -
Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025) -
BanglaByT5: Byte-Level Modelling for Bangla
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2025) -
Vacaspati: A Diverse Corpus of Bangla Literature
von: Bhattacharyya, Pramit, et al.
Veröffentlicht: (2023) -
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages
von: Niyogi, Mitodru, et al.
Veröffentlicht: (2024)