Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yulong, Schlegel, Viktor, Batista-Navarro, Riza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension
by: Schlegel, Viktor, et al.
Published: (2024)
by: Schlegel, Viktor, et al.
Published: (2024)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024)
by: Feher, Darius, et al.
Published: (2024)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
by: Madusanka, Tharindu, et al.
Published: (2025)
by: Madusanka, Tharindu, et al.
Published: (2025)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models Using Synthetic Back-Translation Data
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
Pay Attention to What Matters
by: Silva, Pedro Luiz, et al.
Published: (2024)
by: Silva, Pedro Luiz, et al.
Published: (2024)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
by: Ma, Shengkun, et al.
Published: (2025)
by: Ma, Shengkun, et al.
Published: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
by: Wu, Qinzhuo, et al.
Published: (2026)
by: Wu, Qinzhuo, et al.
Published: (2026)
Multilingual Multi-Aspect Explainability Analyses on Machine Reading Comprehension Models
by: Cui, Yiming, et al.
Published: (2021)
by: Cui, Yiming, et al.
Published: (2021)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
by: Srivatsa, KV Aditya, et al.
Published: (2025)
by: Srivatsa, KV Aditya, et al.
Published: (2025)
How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
by: Kim, Doyoung, et al.
Published: (2026)
by: Kim, Doyoung, et al.
Published: (2026)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
by: Chaudhary, Manav, et al.
Published: (2024)
by: Chaudhary, Manav, et al.
Published: (2024)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
by: Xue, Xiaona, et al.
Published: (2026)
by: Xue, Xiaona, et al.
Published: (2026)
Large Language Models in Argument Mining: A Survey
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
Paying More Attention to Source Context: Mitigating Unfaithful Translations from Large Language Model
by: Zhang, Hongbin, et al.
Published: (2024)
by: Zhang, Hongbin, et al.
Published: (2024)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?
by: Bhuiya, Neeladri, et al.
Published: (2024)
by: Bhuiya, Neeladri, et al.
Published: (2024)
C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts
by: Qing, Chenxi, et al.
Published: (2026)
by: Qing, Chenxi, et al.
Published: (2026)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
by: Zhou, Xinyun, et al.
Published: (2025)
by: Zhou, Xinyun, et al.
Published: (2025)
Using Natural Language for Human-Robot Collaboration in the Real World
by: Lindes, Peter, et al.
Published: (2025)
by: Lindes, Peter, et al.
Published: (2025)
Multimodality and Attention Increase Alignment in Natural Language Prediction Between Humans and Computational Models
by: Kewenig, Viktor, et al.
Published: (2023)
by: Kewenig, Viktor, et al.
Published: (2023)
Structured Information Matters: Explainable ICD Coding with Patient-Level Knowledge Graphs
by: Li, Mingyang, et al.
Published: (2025)
by: Li, Mingyang, et al.
Published: (2025)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
by: Zhou, Yongxin, et al.
Published: (2025)
by: Zhou, Yongxin, et al.
Published: (2025)
Evaluating Large Language Models for Real-World Engineering Tasks
by: Heesch, Rene, et al.
Published: (2025)
by: Heesch, Rene, et al.
Published: (2025)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Question Generation for Assessing Early Literacy Reading Comprehension
by: Yang, Xiaocheng, et al.
Published: (2025)
by: Yang, Xiaocheng, et al.
Published: (2025)
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
by: Choi, Junhyuk, et al.
Published: (2025)
by: Choi, Junhyuk, et al.
Published: (2025)
Reading Comprehension using Entity-based Memory Network
by: Wang, Xun, et al.
Published: (2016)
by: Wang, Xun, et al.
Published: (2016)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Aspect-based Sentiment Evaluation of Chess Moves (ASSESS): an NLP-based Method for Evaluating Chess Strategies from Textbooks
by: Alrdahi, Haifa, et al.
Published: (2024)
by: Alrdahi, Haifa, et al.
Published: (2024)
MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
by: Binici, Kuluhan, et al.
Published: (2024)
by: Binici, Kuluhan, et al.
Published: (2024)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
by: Pan, Ruihao, et al.
Published: (2026)
by: Pan, Ruihao, et al.
Published: (2026)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Similar Items
-
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025) -
Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension
by: Schlegel, Viktor, et al.
Published: (2024) -
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024) -
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
by: Madusanka, Tharindu, et al.
Published: (2025) -
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
by: Hong, Kung Yin, et al.
Published: (2024)