Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mondal, Ishani, Stokes, Jack W., Jauhar, Sujay Kumar, Yang, Longqi, Wan, Mengting, Xu, Xiaofeng, Song, Xia, Neville, Jennifer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913729931640832
author Mondal, Ishani
Stokes, Jack W.
Jauhar, Sujay Kumar
Yang, Longqi
Wan, Mengting
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
author_facet Mondal, Ishani
Stokes, Jack W.
Jauhar, Sujay Kumar
Yang, Longqi
Wan, Mengting
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
contents LLMs often fail to meet the specialized needs of distinct user groups due to their one-size-fits-all training paradigm \cite{lucy-etal-2024-one} and there is limited research on what personalization aspects each group expect. To address these limitations, we propose a group-aware personalization framework, Group Preference Alignment (GPA), that identifies context-specific variations in conversational preferences across user groups and then steers LLMs to address those preferences. Our approach consists of two steps: (1) Group-Aware Preference Extraction, where maximally divergent user-group preferences are extracted from real-world conversation logs and distilled into interpretable rubrics, and (2) Tailored Response Generation, which leverages these rubrics through two methods: a) Context-Tuned Inference (GAP-CT), that dynamically adjusts responses via context-dependent prompt instructions, and b) Rubric-Finetuning Inference (GPA-FT), which uses the rubrics to generate contrastive synthetic data for personalization of group-specific models via alignment. Experiments demonstrate that our framework significantly improves alignment of the output with respect to user preferences and outperforms baseline methods, while maintaining robust performance on standard benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations
Mondal, Ishani
Stokes, Jack W.
Jauhar, Sujay Kumar
Yang, Longqi
Wan, Mengting
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
Computation and Language
LLMs often fail to meet the specialized needs of distinct user groups due to their one-size-fits-all training paradigm \cite{lucy-etal-2024-one} and there is limited research on what personalization aspects each group expect. To address these limitations, we propose a group-aware personalization framework, Group Preference Alignment (GPA), that identifies context-specific variations in conversational preferences across user groups and then steers LLMs to address those preferences. Our approach consists of two steps: (1) Group-Aware Preference Extraction, where maximally divergent user-group preferences are extracted from real-world conversation logs and distilled into interpretable rubrics, and (2) Tailored Response Generation, which leverages these rubrics through two methods: a) Context-Tuned Inference (GAP-CT), that dynamically adjusts responses via context-dependent prompt instructions, and b) Rubric-Finetuning Inference (GPA-FT), which uses the rubrics to generate contrastive synthetic data for personalization of group-specific models via alignment. Experiments demonstrate that our framework significantly improves alignment of the output with respect to user preferences and outperforms baseline methods, while maintaining robust performance on standard benchmarks.
title Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations
topic Computation and Language
url https://arxiv.org/abs/2503.08035