Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Pingjun, Chen, Beiduo, Peng, Siyao, de Marneffe, Marie-Catherine, Roth, Benjamin, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
by: Chen, Beiduo, et al.
Published: (2024)
by: Chen, Beiduo, et al.
Published: (2024)
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
by: Chen, Beiduo, et al.
Published: (2025)
by: Chen, Beiduo, et al.
Published: (2025)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
by: Muscato, Benedetta, et al.
Published: (2026)
by: Muscato, Benedetta, et al.
Published: (2026)
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
When Annotators Agree but Labels Disagree: The Projection Problem in Stance Detection
by: Zhang, Bowen
Published: (2026)
by: Zhang, Bowen
Published: (2026)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
by: Cheng, Xinyuan, et al.
Published: (2026)
by: Cheng, Xinyuan, et al.
Published: (2026)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
by: Baan, Joris, et al.
Published: (2024)
by: Baan, Joris, et al.
Published: (2024)
Evaluating Large Language Models for Cross-Lingual Retrieval
by: Zuo, Longfei, et al.
Published: (2025)
by: Zuo, Longfei, et al.
Published: (2025)
MaiBaam Annotation Guidelines
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
by: Casola, Silvia, et al.
Published: (2025)
by: Casola, Silvia, et al.
Published: (2025)
Explanation sensitivity to the randomness of large language models: the case of journalistic text classification
by: Bogaert, Jeremie, et al.
Published: (2024)
by: Bogaert, Jeremie, et al.
Published: (2024)
Revisiting Active Learning under (Human) Label Variation
by: Gruber, Cornelia, et al.
Published: (2025)
by: Gruber, Cornelia, et al.
Published: (2025)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
by: Ruiz, Tomas, et al.
Published: (2025)
by: Ruiz, Tomas, et al.
Published: (2025)
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
Humans and LLMs Diverge on Probabilistic Inferences
by: Kamath, Gaurav, et al.
Published: (2026)
by: Kamath, Gaurav, et al.
Published: (2026)
MultiClimate: Multimodal Stance Detection on Climate Change Videos
by: Wang, Jiawen, et al.
Published: (2024)
by: Wang, Jiawen, et al.
Published: (2024)
Agree to Agree
Published: (2020)
Published: (2020)
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
by: Pei, Renhao, et al.
Published: (2026)
by: Pei, Renhao, et al.
Published: (2026)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
EEVEE: An Easy Annotation Tool for Natural Language Processing
by: Sorensen, Axel, et al.
Published: (2024)
by: Sorensen, Axel, et al.
Published: (2024)
Explanation Regularisation through the Lens of Attributions
by: Ferreira, Pedro, et al.
Published: (2024)
by: Ferreira, Pedro, et al.
Published: (2024)
Neural Text Normalization for Luxembourgish using Real-Life Variation Data
by: Lutgen, Anne-Marie, et al.
Published: (2024)
by: Lutgen, Anne-Marie, et al.
Published: (2024)
A survey of diversity quantification in natural language processing: The why, what, where and how
by: Estève, Louis, et al.
Published: (2025)
by: Estève, Louis, et al.
Published: (2025)
Variation is the Norm: Embracing Sociolinguistics in NLP
by: Lutgen, Anne-Marie, et al.
Published: (2026)
by: Lutgen, Anne-Marie, et al.
Published: (2026)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Agree to Disagree: Measuring Hidden Dissent in FOMC Meetings
by: Tsang, Kwok Ping, et al.
Published: (2023)
by: Tsang, Kwok Ping, et al.
Published: (2023)
Agree to Disagree: Consensus-Free Flocking under Constraints
by: Jardine, Peter Travis, et al.
Published: (2026)
by: Jardine, Peter Travis, et al.
Published: (2026)
Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification
by: Namazov, Mahammad, et al.
Published: (2026)
by: Namazov, Mahammad, et al.
Published: (2026)
Similar Items
-
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
by: Hong, Pingjun, et al.
Published: (2025) -
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024) -
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI
by: Chen, Beiduo, et al.
Published: (2024) -
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026) -
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
by: Chen, Beiduo, et al.
Published: (2024)