IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
Fuente:
arXiv
Guardado en:
| Autores principales: | Khan, Mohammed Safi Ur Rahman, Mehta, Priyam, Sankar, Ananth, Kumaravelan, Umashankar, Doddapaneni, Sumanth, B, Suriyaprasaad, G, Varun Balan, Jain, Sparsh, Kunchukuttan, Anoop, Kumar, Pratyush, Dabre, Raj, Khapra, Mitesh M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
por: Jayakumar, Thanmay, et al.
Publicado: (2026)
por: Jayakumar, Thanmay, et al.
Publicado: (2026)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
por: Sankar, Ashwin, et al.
Publicado: (2024)
por: Sankar, Ashwin, et al.
Publicado: (2024)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
por: Suryanarayanan, Sanjay, et al.
Publicado: (2024)
por: Suryanarayanan, Sanjay, et al.
Publicado: (2024)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
por: Prasanjith, Pasunuti, et al.
Publicado: (2025)
por: Prasanjith, Pasunuti, et al.
Publicado: (2025)
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
por: Singh, Anushka, et al.
Publicado: (2024)
por: Singh, Anushka, et al.
Publicado: (2024)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
por: Mundra, Nandini, et al.
Publicado: (2024)
por: Mundra, Nandini, et al.
Publicado: (2024)
Airavata: Introducing Hindi Instruction-tuned LLM
por: Gala, Jay, et al.
Publicado: (2024)
por: Gala, Jay, et al.
Publicado: (2024)
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
por: Nath, Oikantik, et al.
Publicado: (2025)
por: Nath, Oikantik, et al.
Publicado: (2025)
Can Vision-Language Models Evaluate Handwritten Math?
por: Nath, Oikantik, et al.
Publicado: (2025)
por: Nath, Oikantik, et al.
Publicado: (2025)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2026)
por: Khan, Mohammed Safi Ur Rahman, et al.
Publicado: (2026)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
por: Pulipaka, Sidharth, et al.
Publicado: (2025)
por: Pulipaka, Sidharth, et al.
Publicado: (2025)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
por: Halder, Deepon, et al.
Publicado: (2025)
por: Halder, Deepon, et al.
Publicado: (2025)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
por: Sankar, Ashwin, et al.
Publicado: (2024)
por: Sankar, Ashwin, et al.
Publicado: (2024)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2024)
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
por: Nawale, Janki Atul, et al.
Publicado: (2025)
por: Nawale, Janki Atul, et al.
Publicado: (2025)
CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages
por: Maurya, Kaushal Kumar, et al.
Publicado: (2023)
por: Maurya, Kaushal Kumar, et al.
Publicado: (2023)
MILU: A Multi-task Indic Language Understanding Benchmark
por: Verma, Sshubam, et al.
Publicado: (2024)
por: Verma, Sshubam, et al.
Publicado: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
por: Anand, Srija, et al.
Publicado: (2024)
por: Anand, Srija, et al.
Publicado: (2024)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
por: Husain, Jaavid Aktar, et al.
Publicado: (2024)
por: Husain, Jaavid Aktar, et al.
Publicado: (2024)
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
por: Ghosh, Poulami, et al.
Publicado: (2024)
por: Ghosh, Poulami, et al.
Publicado: (2024)
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
por: Javed, Tahir, et al.
Publicado: (2025)
por: Javed, Tahir, et al.
Publicado: (2025)
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
por: Sankar, Ashwin, et al.
Publicado: (2025)
por: Sankar, Ashwin, et al.
Publicado: (2025)
User Embedding Model for Personalized Language Prompting
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
por: Doddapaneni, Sumanth, et al.
Publicado: (2024)
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
por: Anand, Srija, et al.
Publicado: (2026)
por: Anand, Srija, et al.
Publicado: (2026)
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
por: Javed, Tahir, et al.
Publicado: (2024)
por: Javed, Tahir, et al.
Publicado: (2024)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
por: Anand, Srija, et al.
Publicado: (2024)
por: Anand, Srija, et al.
Publicado: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
por: Bhogale, Kaushal Santosh, et al.
Publicado: (2024)
por: Bhogale, Kaushal Santosh, et al.
Publicado: (2024)
Hybrid Approach to Software Fault Prediction Using Particle Swarm Optimization and Fuzzy Time Series
por: Umashankar Samal
Publicado: (2025)
por: Umashankar Samal
Publicado: (2025)
Software Reliability Growth Model Considering Imperfect Debugging and Fault Removal Efficiency
por: Umashankar Samal
Publicado: (2024)
por: Umashankar Samal
Publicado: (2024)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
por: Kartik, Kartik, et al.
Publicado: (2024)
por: Kartik, Kartik, et al.
Publicado: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
por: Varadhan, Praveen Srinivasa, et al.
Publicado: (2025)
A Study on Web Application Vulnerabilities to find an optimal Security Architecture
por: Amuthadevi, C., et al.
Publicado: (2022)
por: Amuthadevi, C., et al.
Publicado: (2022)
NLIP_Lab-IITH Low-Resource MT System for WMT24 Indic MT Shared Task
por: Sahoo, Pramit, et al.
Publicado: (2024)
por: Sahoo, Pramit, et al.
Publicado: (2024)
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
por: Joshi, Sakshi, et al.
Publicado: (2025)
por: Joshi, Sakshi, et al.
Publicado: (2025)
PERSOMA: PERsonalized SOft ProMpt Adapter Architecture for Personalized Language Prompting
por: Hebert, Liam, et al.
Publicado: (2024)
por: Hebert, Liam, et al.
Publicado: (2024)
Ejemplares similares
-
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
por: Doddapaneni, Sumanth, et al.
Publicado: (2024) -
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
por: Jayakumar, Thanmay, et al.
Publicado: (2026) -
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
por: Doddapaneni, Sumanth, et al.
Publicado: (2024) -
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
por: Sankar, Ashwin, et al.
Publicado: (2024) -
Pralekha: Cross-Lingual Document Alignment for Indic Languages
por: Suryanarayanan, Sanjay, et al.
Publicado: (2024)