Quantifying and Predicting Disagreement in Graded Human Ratings
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Leixin, Çöltekin, Çağrı |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modeling Human Perspectives with Socio-Demographic Representations
by: Zhang, Leixin, et al.
Published: (2026)
by: Zhang, Leixin, et al.
Published: (2026)
Tübingen-CL at SemEval-2024 Task 1:Ensemble Learning for Semantic Relatedness Estimation
by: Zhang, Leixin, et al.
Published: (2024)
by: Zhang, Leixin, et al.
Published: (2024)
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies
by: Cekinel, Recep Firat, et al.
Published: (2024)
by: Cekinel, Recep Firat, et al.
Published: (2024)
Cross-Lingual Learning vs. Low-Resource Fine-Tuning: A Case Study with Fact-Checking in Turkish
by: Cekinel, Recep Firat, et al.
Published: (2024)
by: Cekinel, Recep Firat, et al.
Published: (2024)
Chapter Mapping to prosody
by: Güneş, Güliz, et al.
Published: (2019)
by: Güneş, Güliz, et al.
Published: (2019)
Multilingual Power and Ideology Identification in the Parliament: a Reference Dataset and Simple Baselines
by: Çöltekin, Çağrı, et al.
Published: (2024)
by: Çöltekin, Çağrı, et al.
Published: (2024)
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
by: Lu, Junyu, et al.
Published: (2026)
by: Lu, Junyu, et al.
Published: (2026)
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
by: Muscato, Benedetta, et al.
Published: (2025)
by: Muscato, Benedetta, et al.
Published: (2025)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
by: Yang, Leixin, et al.
Published: (2023)
by: Yang, Leixin, et al.
Published: (2023)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
by: Lan, Jian, et al.
Published: (2024)
by: Lan, Jian, et al.
Published: (2024)
A Dataset for Physical and Abstract Plausibility and Sources of Human Disagreement
by: Eichel, Annerose, et al.
Published: (2024)
by: Eichel, Annerose, et al.
Published: (2024)
STEntConv: Predicting Disagreement with Stance Detection and a Signed Graph Convolutional Network
by: Lorge, Isabelle, et al.
Published: (2024)
by: Lorge, Isabelle, et al.
Published: (2024)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
by: Ehara, Yo
Published: (2026)
by: Ehara, Yo
Published: (2026)
Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum
by: Knupleš, Urban, et al.
Published: (2023)
by: Knupleš, Urban, et al.
Published: (2023)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
Leveraging Annotator Disagreement for Text Classification
by: Xu, Jin, et al.
Published: (2024)
by: Xu, Jin, et al.
Published: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language
by: Toraman, Cagri
Published: (2024)
by: Toraman, Cagri
Published: (2024)
NUTMEG: Separating Signal From Noise in Annotator Disagreement
by: Ivey, Jonathan, et al.
Published: (2025)
by: Ivey, Jonathan, et al.
Published: (2025)
Do Differences in Values Influence Disagreements in Online Discussions?
by: van der Meer, Michiel, et al.
Published: (2023)
by: van der Meer, Michiel, et al.
Published: (2023)
From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
by: Jayaweera, Chathuri, et al.
Published: (2025)
by: Jayaweera, Chathuri, et al.
Published: (2025)
LLM-based Automated Grading with Human-in-the-Loop
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
by: Xu, Yinuo, et al.
Published: (2026)
by: Xu, Yinuo, et al.
Published: (2026)
Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems
by: Tajik, Elham, et al.
Published: (2026)
by: Tajik, Elham, et al.
Published: (2026)
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
by: Yang, Eddie, et al.
Published: (2026)
by: Yang, Eddie, et al.
Published: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
by: Xu, Yinuo, et al.
Published: (2025)
by: Xu, Yinuo, et al.
Published: (2025)
Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection
by: Srikanth, Maya, et al.
Published: (2025)
by: Srikanth, Maya, et al.
Published: (2025)
The Gray Area: Characterizing Moderator Disagreement on Reddit
by: Alipour, Shayan, et al.
Published: (2026)
by: Alipour, Shayan, et al.
Published: (2026)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024)
by: Khurana, Urja, et al.
Published: (2024)
A Decomposition-Based Approach for Evaluating and Analyzing Inter-Annotator Disagreement
by: Levi, Effi, et al.
Published: (2022)
by: Levi, Effi, et al.
Published: (2022)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
Entropy, Disagreement, and the Limits of Foundation Models in Genomics
by: Rochkoulets, Maxime, et al.
Published: (2026)
by: Rochkoulets, Maxime, et al.
Published: (2026)
From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification
by: Xu, Shanshan, et al.
Published: (2023)
by: Xu, Shanshan, et al.
Published: (2023)
Quantifying the Role of Textual Predictability in Automatic Speech Recognition
by: Robertson, Sean, et al.
Published: (2024)
by: Robertson, Sean, et al.
Published: (2024)
LLMs Do Not Grade Essays Like Humans
by: Mathew, Jerin George, et al.
Published: (2026)
by: Mathew, Jerin George, et al.
Published: (2026)
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
by: Huynh, Jessica, et al.
Published: (2026)
by: Huynh, Jessica, et al.
Published: (2026)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
by: Okada, Kensuke, et al.
Published: (2026)
by: Okada, Kensuke, et al.
Published: (2026)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
by: Kamp, Jonathan, et al.
Published: (2024)
by: Kamp, Jonathan, et al.
Published: (2024)
Similar Items
-
Modeling Human Perspectives with Socio-Demographic Representations
by: Zhang, Leixin, et al.
Published: (2026) -
Tübingen-CL at SemEval-2024 Task 1:Ensemble Learning for Semantic Relatedness Estimation
by: Zhang, Leixin, et al.
Published: (2024) -
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies
by: Cekinel, Recep Firat, et al.
Published: (2024) -
Cross-Lingual Learning vs. Low-Resource Fine-Tuning: A Case Study with Fact-Checking in Turkish
by: Cekinel, Recep Firat, et al.
Published: (2024) -
Chapter Mapping to prosody
by: Güneş, Güliz, et al.
Published: (2019)