ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Banerjee, Somnath, Layek, Sayan, Adak, Sayantan, Pechenizkiy, Mykola, Mukherjee, Animesh, Hazra, Rima
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918235836776448
author Banerjee, Somnath
Layek, Sayan
Adak, Sayantan
Pechenizkiy, Mykola
Mukherjee, Animesh
Hazra, Rima
author_facet Banerjee, Somnath
Layek, Sayan
Adak, Sayantan
Pechenizkiy, Mykola
Mukherjee, Animesh
Hazra, Rima
contents Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can amplify risk. We propose ProSocialAlign, a test-time, parameter-efficient framework that steers generation toward safe, empathetic, and value-aligned responses without retraining the base model. We formalize five human-centered objectives and cast safety as lexicographic constrained generation: first, applying hard constraints to eliminate harmful continuations; then optimizing for prosocial quality within the safe set. Our method combines (i) directional regulation, a harm-mitigation mechanism that subtracts a learned "harm vector" in parameter space, and (ii) preference-aware autoregressive reward modeling trained jointly across attributes with gradient conflict resolution, enabling fine-grained, user-controllable decoding. Empirical evaluations across five safety benchmarks demonstrate state-of-the-art performance, reducing unsafe leakage and boosting alignment to human values, with strong gains across multiple evaluation metrics. ProSocialAlign offers a robust and modular foundation for generating context-sensitive, safe, and human-aligned responses at inference time.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
Banerjee, Somnath
Layek, Sayan
Adak, Sayantan
Pechenizkiy, Mykola
Mukherjee, Animesh
Hazra, Rima
Computation and Language
Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can amplify risk. We propose ProSocialAlign, a test-time, parameter-efficient framework that steers generation toward safe, empathetic, and value-aligned responses without retraining the base model. We formalize five human-centered objectives and cast safety as lexicographic constrained generation: first, applying hard constraints to eliminate harmful continuations; then optimizing for prosocial quality within the safe set. Our method combines (i) directional regulation, a harm-mitigation mechanism that subtracts a learned "harm vector" in parameter space, and (ii) preference-aware autoregressive reward modeling trained jointly across attributes with gradient conflict resolution, enabling fine-grained, user-controllable decoding. Empirical evaluations across five safety benchmarks demonstrate state-of-the-art performance, reducing unsafe leakage and boosting alignment to human values, with strong gains across multiple evaluation metrics. ProSocialAlign offers a robust and modular foundation for generating context-sensitive, safe, and human-aligned responses at inference time.
title ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
topic Computation and Language
url https://arxiv.org/abs/2512.06515