Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Jingwei, Fan, Yu, Zouhar, Vilém, Rooein, Donya, Hoyle, Alexander, Sachan, Mrinmaya, Leippold, Markus, Hovy, Dirk, Ash, Elliott |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
by: Xiong, Chenfei, et al.
Published: (2025)
by: Xiong, Chenfei, et al.
Published: (2025)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
Multilingual Performance Biases of Large Language Models in Education
by: Gupta, Vansh, et al.
Published: (2025)
by: Gupta, Vansh, et al.
Published: (2025)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting
by: Chowdhury, Sankalan Pal, et al.
Published: (2025)
by: Chowdhury, Sankalan Pal, et al.
Published: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
PATS: Personality-Aware Teaching Strategies with Large Language Model Tutors
by: Rooein, Donya, et al.
Published: (2026)
by: Rooein, Donya, et al.
Published: (2026)
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
by: Zengaffinen, Yanick, et al.
Published: (2026)
by: Zengaffinen, Yanick, et al.
Published: (2026)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
Early-Exit and Instant Confidence Translation Quality Estimation
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
by: Schimanski, Tobias, et al.
Published: (2026)
by: Schimanski, Tobias, et al.
Published: (2026)
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026)
by: Zouhar, Vilém, et al.
Published: (2026)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
by: Züfle, Maike, et al.
Published: (2025)
by: Züfle, Maike, et al.
Published: (2025)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
A Formal Perspective on Byte-Pair Encoding
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical Argumentation
by: Shi, Minjing, et al.
Published: (2026)
by: Shi, Minjing, et al.
Published: (2026)
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
by: Kohler, Benjamin, et al.
Published: (2026)
by: Kohler, Benjamin, et al.
Published: (2026)
ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Multimodal Shannon Game with Images
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
by: Cui, Peng, et al.
Published: (2025)
by: Cui, Peng, et al.
Published: (2025)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
by: Orlikowski, Matthias, et al.
Published: (2023)
by: Orlikowski, Matthias, et al.
Published: (2023)
Similar Items
-
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
by: Xiong, Chenfei, et al.
Published: (2025) -
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
by: Rooein, Donya, et al.
Published: (2025) -
Multilingual Performance Biases of Large Language Models in Education
by: Gupta, Vansh, et al.
Published: (2025) -
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024) -
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
by: Ni, Jingwei, et al.
Published: (2024)