When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Fleisig, Eve, Abebe, Rediet, Klein, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
by: Verma, Vivek, et al.
Published: (2023)
by: Verma, Vivek, et al.
Published: (2023)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025)
by: Fleisig, Eve, et al.
Published: (2025)
Lawma: The Power of Specialization for Legal Annotation
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Mapping Social Choice Theory to RLHF
by: Dai, Jessica, et al.
Published: (2024)
by: Dai, Jessica, et al.
Published: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
by: Fu, Tairan, et al.
Published: (2025)
by: Fu, Tairan, et al.
Published: (2025)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
by: Jaggi, Harbani, et al.
Published: (2024)
by: Jaggi, Harbani, et al.
Published: (2024)
Perspective Transition of Large Language Models for Solving Subjective Tasks
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
by: Dehghan, Somaiyeh, et al.
Published: (2025)
by: Dehghan, Somaiyeh, et al.
Published: (2025)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024)
by: Khurana, Urja, et al.
Published: (2024)
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
by: Ruiz, Tomas, et al.
Published: (2025)
by: Ruiz, Tomas, et al.
Published: (2025)
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
by: Törnberg, Petter
Published: (2026)
by: Törnberg, Petter
Published: (2026)
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
by: Zhao, Justin, et al.
Published: (2024)
by: Zhao, Justin, et al.
Published: (2024)
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
GPTs Are Multilingual Annotators for Sequence Generation Tasks
by: Choi, Juhwan, et al.
Published: (2024)
by: Choi, Juhwan, et al.
Published: (2024)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
by: Munir, Sheza, et al.
Published: (2026)
by: Munir, Sheza, et al.
Published: (2026)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
by: Kamp, Jonathan, et al.
Published: (2024)
by: Kamp, Jonathan, et al.
Published: (2024)
FuocChuVIP123 at CoMeDi Shared Task: Disagreement Ranking with XLM-Roberta Sentence Embeddings and Deep Neural Regression
by: Chu, Phuoc Duong Huy
Published: (2025)
by: Chu, Phuoc Duong Huy
Published: (2025)
Disagreements in Reasoning: How a Model's Thinking Process Dictates Persuasion in Multi-Agent Systems
by: Zhao, Haodong, et al.
Published: (2025)
by: Zhao, Haodong, et al.
Published: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
by: Kim, Siun, et al.
Published: (2026)
by: Kim, Siun, et al.
Published: (2026)
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
by: Prabhakaran, Vinodkumar, et al.
Published: (2023)
by: Prabhakaran, Vinodkumar, et al.
Published: (2023)
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
by: Allen, Bradley P., et al.
Published: (2025)
by: Allen, Bradley P., et al.
Published: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
by: Hagar, Nick, et al.
Published: (2025)
by: Hagar, Nick, et al.
Published: (2025)
What's Wrong? Refining Meeting Summaries with LLM Feedback
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
HYBRINFOX at CheckThat! 2024 -- Task 2: Enriching BERT Models with the Expert System VAGO for Subjectivity Detection
by: Casanova, Morgane, et al.
Published: (2024)
by: Casanova, Morgane, et al.
Published: (2024)
American Sign Language Handshapes Reflect Pressures for Communicative Efficiency
by: Yin, Kayo, et al.
Published: (2024)
by: Yin, Kayo, et al.
Published: (2024)
THOUGHTSCULPT: Reasoning with Intermediate Revision and Search
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
by: Arunasalam, Arjun, et al.
Published: (2025)
by: Arunasalam, Arjun, et al.
Published: (2025)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
by: Yang, Kevin, et al.
Published: (2023)
by: Yang, Kevin, et al.
Published: (2023)
Similar Items
-
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
by: Verma, Vivek, et al.
Published: (2023) -
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
by: Fleisig, Eve, et al.
Published: (2025) -
Lawma: The Power of Specialization for Legal Annotation
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024) -
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026) -
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)