Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Banayeeanzade, Amin, Tak, Ala N., Bahrani, Fatemeh, Bolourani, Anahita, Blas, Leonardo, Ferrara, Emilio, Gratch, Jonathan, Karimireddy, Sai Praneeth
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908577384366080
author Banayeeanzade, Amin
Tak, Ala N.
Bahrani, Fatemeh
Bolourani, Anahita
Blas, Leonardo
Ferrara, Emilio
Gratch, Jonathan
Karimireddy, Sai Praneeth
author_facet Banayeeanzade, Amin
Tak, Ala N.
Bahrani, Fatemeh
Bolourani, Anahita
Blas, Leonardo
Ferrara, Emilio
Gratch, Jonathan
Karimireddy, Sai Praneeth
contents The ability to control LLMs' emulated emotional states and personality traits is essential for enabling rich, human-centered interactions in socially interactive settings. We introduce PsySET, a Psychologically-informed benchmark to evaluate LLM Steering Effectiveness and Trustworthiness across the emotion and personality domains. Our study spans four models from different LLM families paired with various steering strategies, including prompting, fine-tuning, and representation engineering. Our results indicate that prompting is consistently effective but limited in intensity control, whereas vector injections achieve finer controllability while slightly reducing output quality. Moreover, we explore the trustworthiness of steered LLMs by assessing safety, truthfulness, fairness, and ethics, highlighting potential side effects and behavioral shifts. Notably, we observe idiosyncratic effects; for instance, even a positive emotion like joy can degrade robustness to adversarial factuality, lower privacy awareness, and increase preferential bias. Meanwhile, anger predictably elevates toxicity yet strengthens leakage resistance. Our framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04484
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
Banayeeanzade, Amin
Tak, Ala N.
Bahrani, Fatemeh
Bolourani, Anahita
Blas, Leonardo
Ferrara, Emilio
Gratch, Jonathan
Karimireddy, Sai Praneeth
Computation and Language
Artificial Intelligence
The ability to control LLMs' emulated emotional states and personality traits is essential for enabling rich, human-centered interactions in socially interactive settings. We introduce PsySET, a Psychologically-informed benchmark to evaluate LLM Steering Effectiveness and Trustworthiness across the emotion and personality domains. Our study spans four models from different LLM families paired with various steering strategies, including prompting, fine-tuning, and representation engineering. Our results indicate that prompting is consistently effective but limited in intensity control, whereas vector injections achieve finer controllability while slightly reducing output quality. Moreover, we explore the trustworthiness of steered LLMs by assessing safety, truthfulness, fairness, and ethics, highlighting potential side effects and behavioral shifts. Notably, we observe idiosyncratic effects; for instance, even a positive emotion like joy can degrade robustness to adversarial factuality, lower privacy awareness, and increase preferential bias. Meanwhile, anger predictably elevates toxicity yet strengthens leakage resistance. Our framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications.
title Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.04484