Quantifying and Predicting Disagreement in Graded Human Ratings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Leixin, Çöltekin, Çağrı
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917454196768768
author Zhang, Leixin
Çöltekin, Çağrı
author_facet Zhang, Leixin
Çöltekin, Çağrı
contents It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we investigate annotation variation patterns in graded human ratings for inappropriate languages, including offensive language, hate speech, and toxic language perception. We examine whether the degree of annotation disagreement can be predicted from textual features. We further propose the Opposition Index, a metric that quantifies perspective opposition among annotators on a given item, and investigate the predictability of instances with potentially opposing human opinions. Our results show a moderate positive correlation between estimated and observed annotation variance. We find that two approaches achieve comparable performance in variance prediction: directly predicting the variance value and estimating it from predicted annotation distributions. Our results on opposition perspective prediction show that items with high opposition index values are more difficult to predict and are often underestimated by models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01168
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantifying and Predicting Disagreement in Graded Human Ratings
Zhang, Leixin
Çöltekin, Çağrı
Computation and Language
It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all instances in a given task elicit the same degree of opinion divergence. In this paper, we investigate annotation variation patterns in graded human ratings for inappropriate languages, including offensive language, hate speech, and toxic language perception. We examine whether the degree of annotation disagreement can be predicted from textual features. We further propose the Opposition Index, a metric that quantifies perspective opposition among annotators on a given item, and investigate the predictability of instances with potentially opposing human opinions. Our results show a moderate positive correlation between estimated and observed annotation variance. We find that two approaches achieve comparable performance in variance prediction: directly predicting the variance value and estimating it from predicted annotation distributions. Our results on opposition perspective prediction show that items with high opposition index values are more difficult to predict and are often underestimated by models.
title Quantifying and Predicting Disagreement in Graded Human Ratings
topic Computation and Language
url https://arxiv.org/abs/2605.01168