Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fitz, Stephen, Romero, Peter, Basart, Steven, Chen, Sipeng, Hernandez-Orallo, Jose
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911165294051328
author Fitz, Stephen
Romero, Peter
Basart, Steven
Chen, Sipeng
Hernandez-Orallo, Jose
author_facet Fitz, Stephen
Romero, Peter
Basart, Steven
Chen, Sipeng
Hernandez-Orallo, Jose
contents Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consistent and measurable synthetic personality traits, little is known about how modulating these traits affects model behavior. We address this gap by investigating how psychometric personality control grounded in the Big Five framework influences AI behavior in the context of capability and safety benchmarks. Our experiments reveal striking effects: for example, reducing conscientiousness leads to significant drops in safety-relevant metrics on benchmarks such as WMDP, TruthfulQA, ETHICS, and Sycophancy as well as reduction in general capabilities as measured by MMLU. These findings highlight personality shaping as a powerful and underexplored axis of model control that interacts with both safety and general competence. We discuss the implications for safety evaluation, alignment strategies, steering model behavior after deployment, and risks associated with possible exploitation of these findings. Our findings motivate a new line of research on personality-sensitive safety evaluations and dynamic behavioral control in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
Fitz, Stephen
Romero, Peter
Basart, Steven
Chen, Sipeng
Hernandez-Orallo, Jose
Artificial Intelligence
Computation and Language
Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consistent and measurable synthetic personality traits, little is known about how modulating these traits affects model behavior. We address this gap by investigating how psychometric personality control grounded in the Big Five framework influences AI behavior in the context of capability and safety benchmarks. Our experiments reveal striking effects: for example, reducing conscientiousness leads to significant drops in safety-relevant metrics on benchmarks such as WMDP, TruthfulQA, ETHICS, and Sycophancy as well as reduction in general capabilities as measured by MMLU. These findings highlight personality shaping as a powerful and underexplored axis of model control that interacts with both safety and general competence. We discuss the implications for safety evaluation, alignment strategies, steering model behavior after deployment, and risks associated with possible exploitation of these findings. Our findings motivate a new line of research on personality-sensitive safety evaluations and dynamic behavioral control in LLMs.
title Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.16332