Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Chu Fei, Dahan, Samuel, Zhu, Xiaodan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912657217421312
author Luo, Chu Fei
Dahan, Samuel
Zhu, Xiaodan
author_facet Luo, Chu Fei
Dahan, Samuel
Zhu, Xiaodan
contents As language models have a greater impact on society, it is important to ensure they are aligned to a diverse range of perspectives and are able to reflect nuance in human values. However, the most popular training paradigms for modern language models often assume there is one optimal answer for every query, leading to generic responses and poor alignment. In this work, we aim to enhance pluralistic alignment of language models in a low-resource setting with two methods: pluralistic decoding and model steering. We empirically demonstrate that model steering offers consistent improvement over zero-shot and few-shot baselines with only 50 annotated samples. Our proposed methods decrease false positives in several high-stakes tasks such as hate speech detection and misinformation detection, and improves the distributional alignment to human values in GlobalOpinionQA. We hope our work highlights the importance of diversity and how language models can be adapted to consider nuanced perspectives.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
Luo, Chu Fei
Dahan, Samuel
Zhu, Xiaodan
Computation and Language
As language models have a greater impact on society, it is important to ensure they are aligned to a diverse range of perspectives and are able to reflect nuance in human values. However, the most popular training paradigms for modern language models often assume there is one optimal answer for every query, leading to generic responses and poor alignment. In this work, we aim to enhance pluralistic alignment of language models in a low-resource setting with two methods: pluralistic decoding and model steering. We empirically demonstrate that model steering offers consistent improvement over zero-shot and few-shot baselines with only 50 annotated samples. Our proposed methods decrease false positives in several high-stakes tasks such as hate speech detection and misinformation detection, and improves the distributional alignment to human values in GlobalOpinionQA. We hope our work highlights the importance of diversity and how language models can be adapted to consider nuanced perspectives.
title Towards Low-Resource Alignment to Diverse Perspectives with Sparse Feedback
topic Computation and Language
url https://arxiv.org/abs/2510.16257