Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosh, Poulami, Dabre, Raj, Bhattacharyya, Pushpak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Morphology-Based Investigation of Positional Encodings
by: Ghosh, Poulami, et al.
Published: (2024)
by: Ghosh, Poulami, et al.
Published: (2024)
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
by: Gaikwad, Pranav, et al.
Published: (2024)
by: Gaikwad, Pranav, et al.
Published: (2024)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
by: Prasanjith, Pasunuti, et al.
Published: (2025)
by: Prasanjith, Pasunuti, et al.
Published: (2025)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
by: Yousofi, Waisullah, et al.
Published: (2024)
by: Yousofi, Waisullah, et al.
Published: (2024)
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
by: Halder, Deepon, et al.
Published: (2026)
by: Halder, Deepon, et al.
Published: (2026)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
by: Sravanthi, Settaluri Lakshmi, et al.
Published: (2024)
LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
IndicEval-XL: Bridging Linguistic Diversity in Code Generation Across Indic Languages
by: Singh, Ujjwal, et al.
Published: (2025)
by: Singh, Ujjwal, et al.
Published: (2025)
Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts
by: Oram, Sneha, et al.
Published: (2026)
by: Oram, Sneha, et al.
Published: (2026)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models
by: Jhalani, Manas, et al.
Published: (2024)
by: Jhalani, Manas, et al.
Published: (2024)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
by: Aravapalli, Akhilesh, et al.
Published: (2024)
by: Aravapalli, Akhilesh, et al.
Published: (2024)
Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages
by: Deoghare, Sourabh, et al.
Published: (2024)
by: Deoghare, Sourabh, et al.
Published: (2024)
Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
by: Chitale, Pranjal A., et al.
Published: (2025)
by: Chitale, Pranjal A., et al.
Published: (2025)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
An Empirical Study of In-context Learning in LLMs for Machine Translation
by: Chitale, Pranjal A., et al.
Published: (2024)
by: Chitale, Pranjal A., et al.
Published: (2024)
Improving Text Style Transfer using Masked Diffusion Language Models with Inference-time Scaling
by: Padole, Tejomay Kishor, et al.
Published: (2025)
by: Padole, Tejomay Kishor, et al.
Published: (2025)
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
by: Kumar, Sanjeev, et al.
Published: (2026)
by: Kumar, Sanjeev, et al.
Published: (2026)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
Natural Language Processing for Dialects of a Language: A Survey
by: Joshi, Aditya, et al.
Published: (2024)
by: Joshi, Aditya, et al.
Published: (2024)
A Case Study on Context-Aware Neural Machine Translation with Multi-Task Learning
by: Appicharla, Ramakrishna, et al.
Published: (2024)
by: Appicharla, Ramakrishna, et al.
Published: (2024)
Facts-and-Feelings: Capturing both Objectivity and Subjectivity in Table-to-Text Generation
by: Dey, Tathagata, et al.
Published: (2024)
by: Dey, Tathagata, et al.
Published: (2024)
We Care: Multimodal Depression Detection and Knowledge Infused Mental Health Therapeutic Response Generation
by: Moon, Palash, et al.
Published: (2024)
by: Moon, Palash, et al.
Published: (2024)
P-ReMIS: Pragmatic Reasoning in Mental Health and a Social Implication
by: Oram, Sneha, et al.
Published: (2025)
by: Oram, Sneha, et al.
Published: (2025)
Main Predicate and Their Arguments as Explanation Signals For Intent Classification
by: Pimparkhede, Sameer, et al.
Published: (2025)
by: Pimparkhede, Sameer, et al.
Published: (2025)
The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic Languages
by: Deroy, Aniket
Published: (2026)
by: Deroy, Aniket
Published: (2026)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
by: Kaing, Hour, et al.
Published: (2025)
by: Kaing, Hour, et al.
Published: (2025)
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
by: Kammakomati, Mehant, et al.
Published: (2024)
by: Kammakomati, Mehant, et al.
Published: (2024)
Analysis of Indic Language Capabilities in LLMs
by: Vaidya, Aatman, et al.
Published: (2025)
by: Vaidya, Aatman, et al.
Published: (2025)
Statistical Machine Translation for Indic Languages
by: Das, Sudhansu Bala, et al.
Published: (2023)
by: Das, Sudhansu Bala, et al.
Published: (2023)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
Limited-Resource Adapters Are Regularizers, Not Linguists
by: Fekete, Marcell, et al.
Published: (2025)
by: Fekete, Marcell, et al.
Published: (2025)
Similar Items
-
A Morphology-Based Investigation of Positional Encodings
by: Ghosh, Poulami, et al.
Published: (2024) -
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024) -
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
by: Gaikwad, Pranav, et al.
Published: (2024) -
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026) -
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
by: Prasanjith, Pasunuti, et al.
Published: (2025)