Personality Editing for Language Models through Adjusting Self-Referential Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Seojin, Kim, Yumin, Kim, Byeongjeong, Shin, Donghoon, Lee, Hwanhee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917213106077696
author Hwang, Seojin
Kim, Yumin
Kim, Byeongjeong
Shin, Donghoon
Lee, Hwanhee
author_facet Hwang, Seojin
Kim, Yumin
Kim, Byeongjeong
Shin, Donghoon
Lee, Hwanhee
contents Large Language Models (LLMs) are integral to applications such as conversational agents and content creation, where precise control over a model's personality is essential for maintaining tone, consistency, and user engagement. However, prevailing prompt-based or fine-tuning approaches either lack robustness or demand large-scale training data, making them costly and impractical. In this paper, we present PALETTE (Personality Adjustment by LLM SElf-TargeTed quEries), a novel method for personality editing in LLMs. Our approach introduces adjustment queries, where self-referential statements grounded in psychological constructs are treated analogously to factual knowledge, enabling direct editing of personality-related responses. Unlike fine-tuning, PALETTE requires only 12 editing samples to achieve substantial improvements in personality alignment across personality dimensions. Experimental results from both automatic and human evaluations demonstrate that our method enables more stable and well-balanced personality control in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Personality Editing for Language Models through Adjusting Self-Referential Queries
Hwang, Seojin
Kim, Yumin
Kim, Byeongjeong
Shin, Donghoon
Lee, Hwanhee
Computation and Language
Large Language Models (LLMs) are integral to applications such as conversational agents and content creation, where precise control over a model's personality is essential for maintaining tone, consistency, and user engagement. However, prevailing prompt-based or fine-tuning approaches either lack robustness or demand large-scale training data, making them costly and impractical. In this paper, we present PALETTE (Personality Adjustment by LLM SElf-TargeTed quEries), a novel method for personality editing in LLMs. Our approach introduces adjustment queries, where self-referential statements grounded in psychological constructs are treated analogously to factual knowledge, enabling direct editing of personality-related responses. Unlike fine-tuning, PALETTE requires only 12 editing samples to achieve substantial improvements in personality alignment across personality dimensions. Experimental results from both automatic and human evaluations demonstrate that our method enables more stable and well-balanced personality control in LLMs.
title Personality Editing for Language Models through Adjusting Self-Referential Queries
topic Computation and Language
url https://arxiv.org/abs/2502.11789