Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bhandari, Pranav, Fay, Nicolas, Selvaganapathy, Sanjeevan, Datta, Amitava, Naseem, Usman, Nasim, Mehwish
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915838897946624
author Bhandari, Pranav
Fay, Nicolas
Selvaganapathy, Sanjeevan
Datta, Amitava
Naseem, Usman
Nasim, Mehwish
author_facet Bhandari, Pranav
Fay, Nicolas
Selvaganapathy, Sanjeevan
Datta, Amitava
Naseem, Usman
Nasim, Mehwish
contents Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The need for effective mechanisms for behavioural manipulation of the model during generation is a critical gap in the literature that needs to be fulfilled. Personality-aware LLMs hold a promising direction towards this objective. However, the relationship between these psychological constructs and their representations within LLMs remains underexplored and requires further investigation. Moreover, it is intriguing to understand and study the use of these representations to steer the models' behaviour. We propose a novel pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism), which is a comprehensive and empirically validated framework to model human personality applies low-rank subspace discovery methods, and identifies trait-specific optimal layers across different model architectures for robust injection. The resulting personality-aligned directions are then operationalised through a flexible steering framework with dynamic layer selection, enabling precise control of trait expression in LLM outputs. Our findings reveal that personality traits occupy a low-rank shared subspace, and that these latent structures can be transformed into actionable mechanisms for effective steering through careful perturbations without impacting the fluency, variance and general capabilities, helping to bridge the gap between psychological theory and practical model alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
Bhandari, Pranav
Fay, Nicolas
Selvaganapathy, Sanjeevan
Datta, Amitava
Naseem, Usman
Nasim, Mehwish
Computation and Language
Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The need for effective mechanisms for behavioural manipulation of the model during generation is a critical gap in the literature that needs to be fulfilled. Personality-aware LLMs hold a promising direction towards this objective. However, the relationship between these psychological constructs and their representations within LLMs remains underexplored and requires further investigation. Moreover, it is intriguing to understand and study the use of these representations to steer the models' behaviour. We propose a novel pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism), which is a comprehensive and empirically validated framework to model human personality applies low-rank subspace discovery methods, and identifies trait-specific optimal layers across different model architectures for robust injection. The resulting personality-aligned directions are then operationalised through a flexible steering framework with dynamic layer selection, enabling precise control of trait expression in LLM outputs. Our findings reveal that personality traits occupy a low-rank shared subspace, and that these latent structures can be transformed into actionable mechanisms for effective steering through careful perturbations without impacting the fluency, variance and general capabilities, helping to bridge the gap between psychological theory and practical model alignment.
title Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
topic Computation and Language
url https://arxiv.org/abs/2511.03738