Do LLMs "Feel"? Emotion Circuits Discovery and Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chenxi, Zhang, Yixuan, Yu, Ruiji, Zheng, Yufei, Gao, Lang, Song, Zirui, Xu, Zixiang, Xia, Gus, Zhang, Huishuai, Zhao, Dongyan, Chen, Xiuying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914089863741440
author Wang, Chenxi
Zhang, Yixuan
Yu, Ruiji
Zheng, Yufei
Gao, Lang
Song, Zirui
Xu, Zixiang
Xia, Gus
Zhang, Huishuai
Zhao, Dongyan
Chen, Xiuying
author_facet Wang, Chenxi
Zhang, Yixuan
Yu, Ruiji
Zheng, Yufei
Gao, Lang
Song, Zirui
Xu, Zixiang
Xia, Gus
Zhang, Huishuai
Zhao, Dongyan
Chen, Xiuying
contents As the demand for emotional intelligence in large language models (LLMs) grows, a key challenge lies in understanding the internal mechanisms that give rise to emotional expression and in controlling emotions in generated text. This study addresses three core questions: (1) Do LLMs contain context-agnostic mechanisms shaping emotional expression? (2) What form do these mechanisms take? (3) Can they be harnessed for universal emotion control? We first construct a controlled dataset, SEV (Scenario-Event with Valence), to elicit comparable internal states across emotions. Subsequently, we extract context-agnostic emotion directions that reveal consistent, cross-context encoding of emotion (Q1). We identify neurons and attention heads that locally implement emotional computation through analytical decomposition and causal analysis, and validate their causal roles via ablation and enhancement interventions. Next, we quantify each sublayer's causal influence on the model's final emotion representation and integrate the identified local components into coherent global emotion circuits that drive emotional expression (Q2). Directly modulating these circuits achieves 99.65% emotion-expression accuracy on the test set, surpassing prompting- and steering-based methods (Q3). To our knowledge, this is the first systematic study to uncover and validate emotion circuits in LLMs, offering new insights into interpretability and controllable emotional intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do LLMs "Feel"? Emotion Circuits Discovery and Control
Wang, Chenxi
Zhang, Yixuan
Yu, Ruiji
Zheng, Yufei
Gao, Lang
Song, Zirui
Xu, Zixiang
Xia, Gus
Zhang, Huishuai
Zhao, Dongyan
Chen, Xiuying
Computation and Language
Artificial Intelligence
As the demand for emotional intelligence in large language models (LLMs) grows, a key challenge lies in understanding the internal mechanisms that give rise to emotional expression and in controlling emotions in generated text. This study addresses three core questions: (1) Do LLMs contain context-agnostic mechanisms shaping emotional expression? (2) What form do these mechanisms take? (3) Can they be harnessed for universal emotion control? We first construct a controlled dataset, SEV (Scenario-Event with Valence), to elicit comparable internal states across emotions. Subsequently, we extract context-agnostic emotion directions that reveal consistent, cross-context encoding of emotion (Q1). We identify neurons and attention heads that locally implement emotional computation through analytical decomposition and causal analysis, and validate their causal roles via ablation and enhancement interventions. Next, we quantify each sublayer's causal influence on the model's final emotion representation and integrate the identified local components into coherent global emotion circuits that drive emotional expression (Q2). Directly modulating these circuits achieves 99.65% emotion-expression accuracy on the test set, surpassing prompting- and steering-based methods (Q3). To our knowledge, this is the first systematic study to uncover and validate emotion circuits in LLMs, offering new insights into interpretability and controllable emotional intelligence.
title Do LLMs "Feel"? Emotion Circuits Discovery and Control
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.11328