Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
Fuente:
arXiv
Saved in:
| Main Authors: | Yi, Zihao, Jiang, Qingxuan, Ma, Ruotian, Chen, Xingyu, Yang, Qu, Wang, Mengru, Ye, Fanghua, Shen, Ying, Tu, Zhaopeng, Li, Xiaolong, Linus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
by: Ma, Xinbei, et al.
Published: (2025)
by: Ma, Xinbei, et al.
Published: (2025)
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
by: Zhang, Bang, et al.
Published: (2025)
by: Zhang, Bang, et al.
Published: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
by: Shi, Zhengliang, et al.
Published: (2025)
by: Shi, Zhengliang, et al.
Published: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
by: Wang, Peisong, et al.
Published: (2025)
by: Wang, Peisong, et al.
Published: (2025)
Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents
by: Yang, Ruihan, et al.
Published: (2026)
by: Yang, Ruihan, et al.
Published: (2026)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
Benchmarking LLMs via Uncertainty Quantification
by: Ye, Fanghua, et al.
Published: (2024)
by: Ye, Fanghua, et al.
Published: (2024)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
by: Wang, Mengru, et al.
Published: (2025)
by: Wang, Mengru, et al.
Published: (2025)
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
by: Yuan, Youliang, et al.
Published: (2023)
by: Yuan, Youliang, et al.
Published: (2023)
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
by: Wang, Mengru, et al.
Published: (2025)
by: Wang, Mengru, et al.
Published: (2025)
No Villains
by: Berry, John N., III
Published: (2009)
by: Berry, John N., III
Published: (2009)
"Good" and "Bad" Failures in Industrial CI/CD -- Balancing Cost and Quality Assurance
by: Sun, Simin, et al.
Published: (2025)
by: Sun, Simin, et al.
Published: (2025)
Leaders or Villains? The Role of Corruption in Shaping the Stereotypes of Politicians
by: Inês Ascenso, et al.
Published: (2025)
by: Inês Ascenso, et al.
Published: (2025)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
"Global is Good, Local is Bad?": Understanding Brand Bias in LLMs
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
by: Kamruzzaman, Mahammed, et al.
Published: (2024)
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
by: Pang, Jianhui, et al.
Published: (2024)
by: Pang, Jianhui, et al.
Published: (2024)
Take the Good With the Bad, and the Bad With the Good? An Experiment on Pro‐Environmental Compensatory Behavior
by: Sophie Clot, et al.
Published: (2025)
by: Sophie Clot, et al.
Published: (2025)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
From LLMs to LRMs: Rethinking Pruning for Reasoning-Centric Models
by: Ding, Longwei, et al.
Published: (2026)
by: Ding, Longwei, et al.
Published: (2026)
Be Transparent in Good Times and Bad
Published: (2024)
Published: (2024)
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
Good Allocations from Bad Estimates
by: Casacuberta, Sílvia, et al.
Published: (2026)
by: Casacuberta, Sílvia, et al.
Published: (2026)
Responsible AI: The Good, The Bad, The AI
by: Jafari, Akbar Anbar, et al.
Published: (2026)
by: Jafari, Akbar Anbar, et al.
Published: (2026)
The Good, the Bad and the Dead! Using Encyclopedias.
by: Bailey, Annette
Published: (1998)
by: Bailey, Annette
Published: (1998)
Wealth Accumulation: The Good, the Bad and the Ugly
by: Stewart Lansley
Published: (2024)
by: Stewart Lansley
Published: (2024)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
From Sensation to Anxiety: The Mediating Effect of Physical Sensation and Experiential Avoidance on Exercise Anxiety Among College Students
by: Yang Qingxuan, et al.
Published: (2026)
by: Yang Qingxuan, et al.
Published: (2026)
LLMs are Good Action Recognizers
by: Qu, Haoxuan, et al.
Published: (2024)
by: Qu, Haoxuan, et al.
Published: (2024)
Villain action in lattice gauge theory
by: Chevyrev, Ilya, et al.
Published: (2024)
by: Chevyrev, Ilya, et al.
Published: (2024)
Parallel spin wave for the Villain model
by: Dario, Paul, et al.
Published: (2025)
by: Dario, Paul, et al.
Published: (2025)
Library Vandalism and the Physical Education Villains
by: Turner, Edward T., et al.
Published: (1973)
by: Turner, Edward T., et al.
Published: (1973)
Are Large Language Models Good Prompt Optimizers?
by: Ma, Ruotian, et al.
Published: (2024)
by: Ma, Ruotian, et al.
Published: (2024)
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
Similar Items
-
BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
by: Wang, Yue, et al.
Published: (2025) -
The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
by: Ma, Xinbei, et al.
Published: (2025) -
Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
by: Zhang, Bang, et al.
Published: (2025) -
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
by: Shi, Zhengliang, et al.
Published: (2025) -
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
by: Liu, Cheng, et al.
Published: (2025)