Airavata: Introducing Hindi Instruction-tuned LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Gala, Jay, Jayakumar, Thanmay, Husain, Jaavid Aktar, M, Aswanth Kumar, Khan, Mohammed Safi Ur Rahman, Kanojia, Diptesh, Puduppully, Ratish, Khapra, Mitesh M., Dabre, Raj, Murthy, Rudra, Kunchukuttan, Anoop |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
by: Husain, Jaavid Aktar, et al.
Published: (2024)
by: Husain, Jaavid Aktar, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
by: Singh, Anushka, et al.
Published: (2024)
by: Singh, Anushka, et al.
Published: (2024)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
by: Mundra, Nandini, et al.
Published: (2024)
by: Mundra, Nandini, et al.
Published: (2024)
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
by: Sankar, Ashwin, et al.
Published: (2024)
by: Sankar, Ashwin, et al.
Published: (2024)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
by: Husain, Jaavid Aktar, et al.
Published: (2026)
by: Husain, Jaavid Aktar, et al.
Published: (2026)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
by: Prasanjith, Pasunuti, et al.
Published: (2025)
by: Prasanjith, Pasunuti, et al.
Published: (2025)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
The Illusion of Generalization in Tabular Language Models
by: Gorla, Aditya, et al.
Published: (2026)
by: Gorla, Aditya, et al.
Published: (2026)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
Can Vision-Language Models Evaluate Handwritten Math?
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Improving Genomic Models via Task-Specific Self-Pretraining
by: Mupparapu, Sohan, et al.
Published: (2025)
by: Mupparapu, Sohan, et al.
Published: (2025)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
Natural Language Processing for Dialects of a Language: A Survey
by: Joshi, Aditya, et al.
Published: (2024)
by: Joshi, Aditya, et al.
Published: (2024)
An Empirical Study of In-context Learning in LLMs for Machine Translation
by: Chitale, Pranjal A., et al.
Published: (2024)
by: Chitale, Pranjal A., et al.
Published: (2024)
Chimera: State Space Models Beyond Sequences
by: Lahoti, Aakash, et al.
Published: (2025)
by: Lahoti, Aakash, et al.
Published: (2025)
Edit Distances and Their Applications to Downstream Tasks in Research and Commercial Contexts
by: Carmo, Félix do, et al.
Published: (2024)
by: Carmo, Félix do, et al.
Published: (2024)
LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
by: Javed, Tahir, et al.
Published: (2024)
by: Javed, Tahir, et al.
Published: (2024)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
by: Han, Vernon Toh Yan, et al.
Published: (2023)
by: Han, Vernon Toh Yan, et al.
Published: (2023)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
by: Nawale, Janki Atul, et al.
Published: (2025)
by: Nawale, Janki Atul, et al.
Published: (2025)
Parameter-Efficient Quality Estimation via Frozen Recursive Models
by: Abubacar, Umar, et al.
Published: (2026)
by: Abubacar, Umar, et al.
Published: (2026)
Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages
by: Deoghare, Sourabh, et al.
Published: (2024)
by: Deoghare, Sourabh, et al.
Published: (2024)
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing
by: Deoghare, Sourabh, et al.
Published: (2025)
by: Deoghare, Sourabh, et al.
Published: (2025)
MILU: A Multi-task Indic Language Understanding Benchmark
by: Verma, Sshubam, et al.
Published: (2024)
by: Verma, Sshubam, et al.
Published: (2024)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
by: Koushik, Girish A., et al.
Published: (2025)
by: Koushik, Girish A., et al.
Published: (2025)
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
by: Javed, Tahir, et al.
Published: (2025)
by: Javed, Tahir, et al.
Published: (2025)
Leveraging Linguistically Enhanced Embeddings for Open Information Extraction
by: Farooqui, Fauzan, et al.
Published: (2024)
by: Farooqui, Fauzan, et al.
Published: (2024)
BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English
by: Srirag, Dipankar, et al.
Published: (2024)
by: Srirag, Dipankar, et al.
Published: (2024)
DGFM: Full Body Dance Generation Driven by Music Foundation Models
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Experiences from Creating a Benchmark for Sentiment Classification for Varieties of English
by: Srirag, Dipankar, et al.
Published: (2024)
by: Srirag, Dipankar, et al.
Published: (2024)
Similar Items
-
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
by: Husain, Jaavid Aktar, et al.
Published: (2024) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025) -
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026) -
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025) -
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
by: Singh, Anushka, et al.
Published: (2024)