Steerable Chatbots: Personalizing LLMs with Preference-Based Activation Steering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bo, Jessica Y., Xu, Tianyu, Chatterjee, Ishan, Passarella-Ward, Katrina, Kulshrestha, Achin, Shin, D
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908362394828800
author Bo, Jessica Y.
Xu, Tianyu
Chatterjee, Ishan
Passarella-Ward, Katrina
Kulshrestha, Achin
Shin, D
author_facet Bo, Jessica Y.
Xu, Tianyu
Chatterjee, Ishan
Passarella-Ward, Katrina
Kulshrestha, Achin
Shin, D
contents As large language models (LLMs) improve in their capacity to serve as personal AI assistants, their ability to output uniquely tailored, personalized responses that align with the soft preferences of their users is essential for enhancing user satisfaction and retention. However, untrained lay users have poor prompt specification abilities and often struggle with conveying their latent preferences to AI assistants. To address this, we leverage activation steering to guide LLMs to align with interpretable preference dimensions during inference. In contrast to memory-based personalization methods that require longer user history, steering is extremely lightweight and can be easily controlled by the user via an linear strength factor. We embed steering into three different interactive chatbot interfaces and conduct a within-subjects user study (n=14) to investigate how end users prefer to personalize their conversations. The results demonstrate the effectiveness of preference-based steering for aligning real-world conversations with hidden user preferences, and highlight further insights on how diverse values around control, usability, and transparency lead users to prefer different interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04260
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steerable Chatbots: Personalizing LLMs with Preference-Based Activation Steering
Bo, Jessica Y.
Xu, Tianyu
Chatterjee, Ishan
Passarella-Ward, Katrina
Kulshrestha, Achin
Shin, D
Human-Computer Interaction
Artificial Intelligence
As large language models (LLMs) improve in their capacity to serve as personal AI assistants, their ability to output uniquely tailored, personalized responses that align with the soft preferences of their users is essential for enhancing user satisfaction and retention. However, untrained lay users have poor prompt specification abilities and often struggle with conveying their latent preferences to AI assistants. To address this, we leverage activation steering to guide LLMs to align with interpretable preference dimensions during inference. In contrast to memory-based personalization methods that require longer user history, steering is extremely lightweight and can be easily controlled by the user via an linear strength factor. We embed steering into three different interactive chatbot interfaces and conduct a within-subjects user study (n=14) to investigate how end users prefer to personalize their conversations. The results demonstrate the effectiveness of preference-based steering for aligning real-world conversations with hidden user preferences, and highlight further insights on how diverse values around control, usability, and transparency lead users to prefer different interfaces.
title Steerable Chatbots: Personalizing LLMs with Preference-Based Activation Steering
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2505.04260