Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Muscato, Benedetta, Chen, Beiduo, Gezici, Gizem, Plank, Barbara, Giannotti, Fosca
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918531762749440
author Muscato, Benedetta
Chen, Beiduo
Gezici, Gizem
Plank, Barbara
Giannotti, Fosca
author_facet Muscato, Benedetta
Chen, Beiduo
Gezici, Gizem
Plank, Barbara
Giannotti, Fosca
contents Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains far less explored. At the same time, it is unclear how to best evaluate human labels and rationales -- or even how to best aggregate rationales beyond majority vote -- in light of this variation. Yet, rationales may provide additional insights into the richness of human reasoning, that may differ in style, values and interpretations -- especially in subjective NLP tasks like hate speech detection. In this work, we unify diverse models, training strategies, loss functions, and existing evaluation metrics under a single protocol by systematically re-implementing them across different label and rationale representation spaces. Classification metrics are organized around two key properties -- predictive and distributional -- while explainability metrics through three complementary dimensions: plausibility, faithfulness, and complexity. In this unified supervision framework, we evaluate model behavior across classification and explainability metrics, as well as metric sensitivity to the choice of label (hard and soft) and rationale representation space (hard, intermediate and soft). Results show that both hard and soft metrics favor softer representations, highlighting their effectiveness in capturing variation and the need to rethink evaluation in subjective NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31563
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
Muscato, Benedetta
Chen, Beiduo
Gezici, Gizem
Plank, Barbara
Giannotti, Fosca
Computation and Language
Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains far less explored. At the same time, it is unclear how to best evaluate human labels and rationales -- or even how to best aggregate rationales beyond majority vote -- in light of this variation. Yet, rationales may provide additional insights into the richness of human reasoning, that may differ in style, values and interpretations -- especially in subjective NLP tasks like hate speech detection. In this work, we unify diverse models, training strategies, loss functions, and existing evaluation metrics under a single protocol by systematically re-implementing them across different label and rationale representation spaces. Classification metrics are organized around two key properties -- predictive and distributional -- while explainability metrics through three complementary dimensions: plausibility, faithfulness, and complexity. In this unified supervision framework, we evaluate model behavior across classification and explainability metrics, as well as metric sensitivity to the choice of label (hard and soft) and rationale representation space (hard, intermediate and soft). Results show that both hard and soft metrics favor softer representations, highlighting their effectiveness in capturing variation and the need to rethink evaluation in subjective NLP.
title Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
topic Computation and Language
url https://arxiv.org/abs/2605.31563