Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
Fuente:
arXiv
Saved in:
| Main Authors: | Doddapaneni, Sumanth, Khan, Mohammed Safi Ur Rahman, Verma, Sshubam, Khapra, Mitesh M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
Can Vision-Language Models Evaluate Handwritten Math?
by: Nath, Oikantik, et al.
Published: (2025)
by: Nath, Oikantik, et al.
Published: (2025)
MILU: A Multi-task Indic Language Understanding Benchmark
by: Verma, Sshubam, et al.
Published: (2024)
by: Verma, Sshubam, et al.
Published: (2024)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
by: Nawale, Janki Atul, et al.
Published: (2025)
by: Nawale, Janki Atul, et al.
Published: (2025)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
by: Sankar, Ashwin, et al.
Published: (2024)
by: Sankar, Ashwin, et al.
Published: (2024)
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
by: Anand, Srija, et al.
Published: (2026)
by: Anand, Srija, et al.
Published: (2026)
NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
by: Javed, Tahir, et al.
Published: (2025)
by: Javed, Tahir, et al.
Published: (2025)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
Airavata: Introducing Hindi Instruction-tuned LLM
by: Gala, Jay, et al.
Published: (2024)
by: Gala, Jay, et al.
Published: (2024)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
The Rarity Blind Spot: A Framework for Evaluating Statistical Reasoning in LLMs
by: Maekawa, Seiji, et al.
Published: (2025)
by: Maekawa, Seiji, et al.
Published: (2025)
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
by: Singh, Anushka, et al.
Published: (2024)
by: Singh, Anushka, et al.
Published: (2024)
LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
by: Javed, Tahir, et al.
Published: (2024)
by: Javed, Tahir, et al.
Published: (2024)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
by: Anand, Srija, et al.
Published: (2024)
by: Anand, Srija, et al.
Published: (2024)
User Embedding Model for Personalized Language Prompting
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
by: Basmov, Victoria, et al.
Published: (2023)
by: Basmov, Victoria, et al.
Published: (2023)
The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious
by: Schessl, Ferdinand M.
Published: (2026)
by: Schessl, Ferdinand M.
Published: (2026)
Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
by: Bhogale, Kaushal Santosh, et al.
Published: (2026)
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
by: Sankar, Ashwin, et al.
Published: (2025)
by: Sankar, Ashwin, et al.
Published: (2025)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
by: Mundra, Nandini, et al.
Published: (2024)
by: Mundra, Nandini, et al.
Published: (2024)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
by: Anand, Srija, et al.
Published: (2024)
by: Anand, Srija, et al.
Published: (2024)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
by: Shin, Hyungyu, et al.
Published: (2025)
by: Shin, Hyungyu, et al.
Published: (2025)
The State Of TTS: A Case Study with Human Fooling Rates
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2025)
Temporal Blind Spots in Large Language Models
by: Wallat, Jonas, et al.
Published: (2024)
by: Wallat, Jonas, et al.
Published: (2024)
PERSOMA: PERsonalized SOft ProMpt Adapter Architecture for Personalized Language Prompting
by: Hebert, Liam, et al.
Published: (2024)
by: Hebert, Liam, et al.
Published: (2024)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
by: Dou, Yao, et al.
Published: (2026)
by: Dou, Yao, et al.
Published: (2026)
Linguistic Blind Spots in Clinical Decision Extraction
by: Elgaar, Mohamed, et al.
Published: (2026)
by: Elgaar, Mohamed, et al.
Published: (2026)
PEDAL: Enhancing Greedy Decoding with Large Language Models using Diverse Exemplars
by: Prabhu, Sumanth
Published: (2024)
by: Prabhu, Sumanth
Published: (2024)
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
by: Shu, Bangzhao, et al.
Published: (2025)
by: Shu, Bangzhao, et al.
Published: (2025)
Mitigating Social Bias in English and Urdu Language Models Using PRM-Guided Candidate Selection and Sequential Refinement
by: Khan, Muneeb Ur Raheem
Published: (2025)
by: Khan, Muneeb Ur Raheem
Published: (2025)
Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks
by: Zhang, Chuyifei, et al.
Published: (2026)
by: Zhang, Chuyifei, et al.
Published: (2026)
Exploring Embedding Priors in Prompt-Tuning for Improved Interpretability and Control
by: Sedov, Sergey, et al.
Published: (2024)
by: Sedov, Sergey, et al.
Published: (2024)
Linguistic Blind Spots of Large Language Models
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Similar Items
-
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
by: Doddapaneni, Sumanth, et al.
Published: (2024) -
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026) -
Can Vision-Language Models Evaluate Handwritten Math?
by: Nath, Oikantik, et al.
Published: (2025) -
MILU: A Multi-task Indic Language Understanding Benchmark
by: Verma, Sshubam, et al.
Published: (2024) -
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
by: Nawale, Janki Atul, et al.
Published: (2025)