From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chebrolu, Niranjan, Jaidka, Kokil, Yeo, Gerard Christopher
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911521475395584
author Chebrolu, Niranjan
Jaidka, Kokil
Yeo, Gerard Christopher
author_facet Chebrolu, Niranjan
Jaidka, Kokil
Yeo, Gerard Christopher
contents Complex social behaviors, such as empathy and strategic politeness, are widely assumed to resist the directional decomposition that makes activation steering effective for coarse attributes like sentiment or toxicity. We present STAR: Steering via Attribution and Representation, which tests this assumption by using attribution patching to identify the layer--token positions where each behavioral trait causally originates, then injecting contrastive activation vectors at precisely those locations. Evaluated on emotional dialogue and negotiation in both single- and multi-turn settings, localized injection consistently outperforms global steering and instruction priming; human evaluation confirms that gains reflect genuine improvements in perceived quality rather than lexical surface change. Our results suggest that complex interpersonal behaviors are encoded as localized, approximately linear directions in LLM activation space, and that behavioral alignment is fundamentally a localization problem.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
Chebrolu, Niranjan
Jaidka, Kokil
Yeo, Gerard Christopher
Computation and Language
Artificial Intelligence
Complex social behaviors, such as empathy and strategic politeness, are widely assumed to resist the directional decomposition that makes activation steering effective for coarse attributes like sentiment or toxicity. We present STAR: Steering via Attribution and Representation, which tests this assumption by using attribution patching to identify the layer--token positions where each behavioral trait causally originates, then injecting contrastive activation vectors at precisely those locations. Evaluated on emotional dialogue and negotiation in both single- and multi-turn settings, localized injection consistently outperforms global steering and instruction priming; human evaluation confirms that gains reflect genuine improvements in perceived quality rather than lexical surface change. Our results suggest that complex interpersonal behaviors are encoded as localized, approximately linear directions in LLM activation space, and that behavioral alignment is fundamentally a localization problem.
title From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.12832