RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lai, Huayi, Song, Shichao, Niu, Simin, Wang, Hanyu, Yang, Jiawei, Wang, Zhouxing, Yin, Zhiqiang, Liang, Xun
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917553254694912
author Lai, Huayi
Song, Shichao
Niu, Simin
Wang, Hanyu
Yang, Jiawei
Wang, Zhouxing
Yin, Zhiqiang
Liang, Xun
author_facet Lai, Huayi
Song, Shichao
Niu, Simin
Wang, Hanyu
Yang, Jiawei
Wang, Zhouxing
Yin, Zhiqiang
Liang, Xun
contents Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts. To address this gap, we introduce RoleCDE, the first benchmark designed to evaluate RPAs under structured conflicts between role-specific values and alignment-oriented constraints. RoleCDE formulates role-aware decision making as cognitive dilemma scenarios, jointly evaluating role-scenario grounding, value conflict resolution, and decision tendencies. The benchmark is constructed at scale, covering approximately 8k diverse role profiles and scenarios and nearly 24k dilemma instances across three difficulty levels and eight role categories. Evaluation of several mainstream LLMs reveals a "Role Value Decoupling" phenomenon, where agents systematically default to alignment-and morality-consistent decisions rather than role-specific values when the two conflict, even under explicit role conditioning. This behavior is largely invariant to dilemma difficulty but varies substantially across role categories. We further show that RoleCDE-based fine-tuning effectively mitigates this decoupling by improving value trade-off reasoning, while preserving general role-playing fidelity and general reasoning performance. Code is available at: https://github.com/rabbitrose/RoleCDE.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01552
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents
Lai, Huayi
Song, Shichao
Niu, Simin
Wang, Hanyu
Yang, Jiawei
Wang, Zhouxing
Yin, Zhiqiang
Liang, Xun
Artificial Intelligence
Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts. To address this gap, we introduce RoleCDE, the first benchmark designed to evaluate RPAs under structured conflicts between role-specific values and alignment-oriented constraints. RoleCDE formulates role-aware decision making as cognitive dilemma scenarios, jointly evaluating role-scenario grounding, value conflict resolution, and decision tendencies. The benchmark is constructed at scale, covering approximately 8k diverse role profiles and scenarios and nearly 24k dilemma instances across three difficulty levels and eight role categories. Evaluation of several mainstream LLMs reveals a "Role Value Decoupling" phenomenon, where agents systematically default to alignment-and morality-consistent decisions rather than role-specific values when the two conflict, even under explicit role conditioning. This behavior is largely invariant to dilemma difficulty but varies substantially across role categories. We further show that RoleCDE-based fine-tuning effectively mitigates this decoupling by improving value trade-off reasoning, while preserving general role-playing fidelity and general reasoning performance. Code is available at: https://github.com/rabbitrose/RoleCDE.
title RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2606.01552