EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
Fuente:
arXiv
Saved in:
| Main Authors: | Zuo, Longfei, Plank, Barbara, Peng, Siyao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
by: Wang, Jiawen, et al.
Published: (2024)
by: Wang, Jiawen, et al.
Published: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
MaiBaam Annotation Guidelines
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
by: Ruiz, Tomas, et al.
Published: (2025)
by: Ruiz, Tomas, et al.
Published: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Evaluating Large Language Models for Cross-Lingual Retrieval
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
by: Pei, Renhao, et al.
Published: (2026)
by: Pei, Renhao, et al.
Published: (2026)
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
by: Xu, Ancheng, et al.
Published: (2025)
by: Xu, Ancheng, et al.
Published: (2025)
EEVEE: An Easy Annotation Tool for Natural Language Processing
by: Sorensen, Axel, et al.
Published: (2024)
by: Sorensen, Axel, et al.
Published: (2024)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
by: Casola, Silvia, et al.
Published: (2025)
by: Casola, Silvia, et al.
Published: (2025)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
Exploring Continual Learning of Compositional Generalization in NLI
by: Fu, Xiyan, et al.
Published: (2024)
by: Fu, Xiyan, et al.
Published: (2024)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
by: Bertolazzi, Leonardo, et al.
Published: (2025)
by: Bertolazzi, Leonardo, et al.
Published: (2025)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
by: Jayaweera, Chathuri, et al.
Published: (2025)
by: Jayaweera, Chathuri, et al.
Published: (2025)
For Generated Text, Is NLI-Neutral Text the Best Text?
by: Mersinias, Michail, et al.
Published: (2023)
by: Mersinias, Michail, et al.
Published: (2023)
What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourse
by: Zhou, Shijia, et al.
Published: (2025)
by: Zhou, Shijia, et al.
Published: (2025)
Add Noise, Tasks, or Layers? MaiNLP at the VarDial 2025 Shared Task on Norwegian Dialectal Slot and Intent Detection
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
by: Harary, Sapir, et al.
Published: (2025)
by: Harary, Sapir, et al.
Published: (2025)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study
by: Krückl, Xaver Maria, et al.
Published: (2025)
by: Krückl, Xaver Maria, et al.
Published: (2025)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
Controlling Reading Ease with Gaze-Guided Text Generation
by: Säuberli, Andreas, et al.
Published: (2026)
by: Säuberli, Andreas, et al.
Published: (2026)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Mathematical Reasoning Enhanced LLM for Formula Derivation: A Case Study on Fiber NLI Modellin
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Entailed Between the Lines: Incorporating Implication into NLI
by: Havaldar, Shreya, et al.
Published: (2025)
by: Havaldar, Shreya, et al.
Published: (2025)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Estimating the Causal Effects of Natural Logic Features in Transformer-Based NLI Models
by: Rozanova, Julia, et al.
Published: (2024)
by: Rozanova, Julia, et al.
Published: (2024)
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
by: Ahnert, Georg, et al.
Published: (2025)
by: Ahnert, Georg, et al.
Published: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Similar Items
-
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
by: Chen, Beiduo, et al.
Published: (2024) -
MultiClimate: Multimodal Stance Detection on Climate Change Videos
by: Wang, Jiawen, et al.
Published: (2024) -
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024) -
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
by: Hong, Pingjun, et al.
Published: (2025) -
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
by: Chen, Beiduo, et al.
Published: (2024)