Saved in:
Bibliographic Details
Main Authors: Parappan, Mohammed Fayiz, Henao, Ricardo
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.06631
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908757019066368
author Parappan, Mohammed Fayiz
Henao, Ricardo
author_facet Parappan, Mohammed Fayiz
Henao, Ricardo
contents Building NLP systems for subjective tasks requires one to ensure their alignment to contrasting human values. We propose the MultiCalibrated Subjective Task Learner framework (MC-STL), which clusters annotations into identifiable human value clusters by three approaches (similarity of annotator rationales, expert-value taxonomies or rater's sociocultural descriptors) and calibrates predictions for each value cluster by learning cluster-specific embeddings. We demonstrate MC-STL on several subjective learning settings, including ordinal, binary, and preference learning predictions, and evaluate it on multiple datasets covering toxic chatbot conversations, offensive social media posts, and human preference alignment. The results show that MC-STL consistently outperforms the baselines that ignore the latent value structure of the annotations, delivering gains in discrimination, value-specific calibration, and disagreement-aware metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06631
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Labels have Human Values: Value Calibration of Subjective Tasks
Parappan, Mohammed Fayiz
Henao, Ricardo
Computation and Language
Building NLP systems for subjective tasks requires one to ensure their alignment to contrasting human values. We propose the MultiCalibrated Subjective Task Learner framework (MC-STL), which clusters annotations into identifiable human value clusters by three approaches (similarity of annotator rationales, expert-value taxonomies or rater's sociocultural descriptors) and calibrates predictions for each value cluster by learning cluster-specific embeddings. We demonstrate MC-STL on several subjective learning settings, including ordinal, binary, and preference learning predictions, and evaluate it on multiple datasets covering toxic chatbot conversations, offensive social media posts, and human preference alignment. The results show that MC-STL consistently outperforms the baselines that ignore the latent value structure of the annotations, delivering gains in discrimination, value-specific calibration, and disagreement-aware metrics.
title Labels have Human Values: Value Calibration of Subjective Tasks
topic Computation and Language
url https://arxiv.org/abs/2601.06631