VISPA: Pluralistic Alignment via Automatic Value Selection and Activation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Shenyan, Zhong, Jiayou, Shetty, Anudeex, Ji, Heng, Nakov, Preslav, Naseem, Usman
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911384248254464
author Zheng, Shenyan
Zhong, Jiayou
Shetty, Anudeex
Ji, Heng
Nakov, Preslav
Naseem, Usman
author_facet Zheng, Shenyan
Zhong, Jiayou
Shetty, Anudeex
Ji, Heng
Nakov, Preslav
Naseem, Usman
contents As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspectives. Achieving such pluralism, however, remains challenging. Existing approaches consider limited values or rely on prompt-level interventions, lacking value control and representation. To address this, we introduce VISPA, a training-free pluralistic alignment framework, that enables direct control over value expression by dynamic selection and internal model activation steering. Across extensive empirical studies spanning multiple models and evaluation settings, we show VISPA is performant across all pluralistic alignment modes in healthcare and beyond. Further analysis reveals VISPA is adaptable with different steering initiations, model, and/or values. These results suggest that pluralistic alignment can be achieved through internal activation mechanisms, offering a scalable path toward language models that serves all.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12758
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
Zheng, Shenyan
Zhong, Jiayou
Shetty, Anudeex
Ji, Heng
Nakov, Preslav
Naseem, Usman
Computation and Language
Artificial Intelligence
Machine Learning
As large language models are increasingly used in high-stakes domains, it is essential that their outputs reflect not average} human preference, rather range of varying perspectives. Achieving such pluralism, however, remains challenging. Existing approaches consider limited values or rely on prompt-level interventions, lacking value control and representation. To address this, we introduce VISPA, a training-free pluralistic alignment framework, that enables direct control over value expression by dynamic selection and internal model activation steering. Across extensive empirical studies spanning multiple models and evaluation settings, we show VISPA is performant across all pluralistic alignment modes in healthcare and beyond. Further analysis reveals VISPA is adaptable with different steering initiations, model, and/or values. These results suggest that pluralistic alignment can be achieved through internal activation mechanisms, offering a scalable path toward language models that serves all.
title VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.12758