Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhary, Manav, Gupta, Harshit, Bhat, Savita, Varma, Vasudeva |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BrainStorm @ iREL at #SMM4H 2024: Leveraging Translation and Topical Embeddings for Annotation Detection in Tweets
by: Chaudhary, Manav, et al.
Published: (2024)
by: Chaudhary, Manav, et al.
Published: (2024)
iREL at SemEval-2024 Task 9: Improving Conventional Prompting Methods for Brain Teasers
by: Gupta, Harshit, et al.
Published: (2024)
by: Gupta, Harshit, et al.
Published: (2024)
Graph-Guided Passage Retrieval for Author-Centric Structured Feedback
by: Chitale, Maitreya Prafulla, et al.
Published: (2025)
by: Chitale, Maitreya Prafulla, et al.
Published: (2025)
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
by: Elenjical, Abraham Paul, et al.
Published: (2026)
by: Elenjical, Abraham Paul, et al.
Published: (2026)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
by: Mehta, Rahul, et al.
Published: (2024)
by: Mehta, Rahul, et al.
Published: (2024)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
by: Pan, Ruihao, et al.
Published: (2026)
by: Pan, Ruihao, et al.
Published: (2026)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Towards a Path Dependent Account of Category Fluency
by: Heineman, David, et al.
Published: (2024)
by: Heineman, David, et al.
Published: (2024)
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
by: Zhou, Xinyun, et al.
Published: (2025)
by: Zhou, Xinyun, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
Modeling Understanding of Story-Based Analogies Using Large Language Models
by: Inani, Kalit, et al.
Published: (2025)
by: Inani, Kalit, et al.
Published: (2025)
Spoken Grammar Assessment Using LLM
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
by: Zhou, Yongxin, et al.
Published: (2025)
by: Zhou, Yongxin, et al.
Published: (2025)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
by: Dutta, Avik, et al.
Published: (2025)
by: Dutta, Avik, et al.
Published: (2025)
MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
by: Xue, Qiyao, et al.
Published: (2025)
by: Xue, Qiyao, et al.
Published: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025)
by: Choukrani, Omar, et al.
Published: (2025)
Semantic Operators: A Declarative Model for Rich, AI-based Data Processing
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
by: Chakraborty, Neeloy, et al.
Published: (2025)
by: Chakraborty, Neeloy, et al.
Published: (2025)
Quality Estimation based Feedback Training for Improving Pronoun Translation
by: Dhankhar, Harshit, et al.
Published: (2025)
by: Dhankhar, Harshit, et al.
Published: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
by: Chen, Andong, et al.
Published: (2024)
by: Chen, Andong, et al.
Published: (2024)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
by: Wu, Juntong, et al.
Published: (2025)
by: Wu, Juntong, et al.
Published: (2025)
Are AI-Generated Text Detectors Robust to Adversarial Perturbations?
by: Huang, Guanhua, et al.
Published: (2024)
by: Huang, Guanhua, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
by: Avinash, Karthik, et al.
Published: (2025)
by: Avinash, Karthik, et al.
Published: (2025)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution
by: Zhu, Jizhao, et al.
Published: (2025)
by: Zhu, Jizhao, et al.
Published: (2025)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
by: Wu, Wanxing, et al.
Published: (2026)
by: Wu, Wanxing, et al.
Published: (2026)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
Towards LLM-based Autograding for Short Textual Answers
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
by: Nickel, Christian, et al.
Published: (2026)
by: Nickel, Christian, et al.
Published: (2026)
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Certainty robustness: Evaluating LLM stability under self-challenging prompts
by: Saadat, Mohammadreza, et al.
Published: (2026)
by: Saadat, Mohammadreza, et al.
Published: (2026)
Emotion Classification in Short English Texts using Deep Learning Techniques
by: Bhat, Siddhanth
Published: (2024)
by: Bhat, Siddhanth
Published: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
Similar Items
-
BrainStorm @ iREL at #SMM4H 2024: Leveraging Translation and Topical Embeddings for Annotation Detection in Tweets
by: Chaudhary, Manav, et al.
Published: (2024) -
iREL at SemEval-2024 Task 9: Improving Conventional Prompting Methods for Brain Teasers
by: Gupta, Harshit, et al.
Published: (2024) -
Graph-Guided Passage Retrieval for Author-Centric Structured Feedback
by: Chitale, Maitreya Prafulla, et al.
Published: (2025) -
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
by: Elenjical, Abraham Paul, et al.
Published: (2026) -
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)