Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Yinuo, Derricks, Veronica, Earl, Allison, Jurgens, David
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2508.02853
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909887265505280
author Xu, Yinuo
Derricks, Veronica
Earl, Allison
Jurgens, David
author_facet Xu, Yinuo
Derricks, Veronica
Earl, Allison
Jurgens, David
contents We present an approach to modeling annotator disagreement in subjective NLP tasks through both architectural and data-centric innovations. Our model, DEM-MoE (Demographic-Aware Mixture of Experts), routes inputs to expert subnetworks based on annotator demographics, enabling it to better represent structured, group-level variation compared to prior models. DEM-MoE consistently performs competitively across demographic groups, and shows especially strong results on datasets with high annotator disagreement. To address sparse demographic coverage, we test whether LLM-generated synthetic annotations via zero-shot persona prompting can be used for data imputation. We show these synthetic judgments align moderately well with human annotations on our data and offer a scalable way to potentially enrich training data. We then propose and evaluate approaches for blending real and synthetic data using strategies tailored to dataset structure. We find that the optimal strategies depend on dataset structure. Together, these contributions improve the representation of diverse perspectives.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
Xu, Yinuo
Derricks, Veronica
Earl, Allison
Jurgens, David
Computation and Language
We present an approach to modeling annotator disagreement in subjective NLP tasks through both architectural and data-centric innovations. Our model, DEM-MoE (Demographic-Aware Mixture of Experts), routes inputs to expert subnetworks based on annotator demographics, enabling it to better represent structured, group-level variation compared to prior models. DEM-MoE consistently performs competitively across demographic groups, and shows especially strong results on datasets with high annotator disagreement. To address sparse demographic coverage, we test whether LLM-generated synthetic annotations via zero-shot persona prompting can be used for data imputation. We show these synthetic judgments align moderately well with human annotations on our data and offer a scalable way to potentially enrich training data. We then propose and evaluate approaches for blending real and synthetic data using strategies tailored to dataset structure. We find that the optimal strategies depend on dataset structure. Together, these contributions improve the representation of diverse perspectives.
title Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
topic Computation and Language
url https://arxiv.org/abs/2508.02853