Steering at the Source: Style Modulation Heads for Robust Persona Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Izawa, Yoshihiro, Minegishi, Gouki, Eguchi, Koshi, Hosokawa, Sosuke, Taura, Kenjiro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2025)
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2025)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
von: Rupprecht, Jens, et al.
Veröffentlicht: (2025)
von: Rupprecht, Jens, et al.
Veröffentlicht: (2025)
Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
von: Gelberg, Yoav, et al.
Veröffentlicht: (2025)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2025)
The Need for a Socially-Grounded Persona Framework for User Simulation
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2026)
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2026)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
von: Kim, Jiseon, et al.
Veröffentlicht: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2025)
von: Venkit, Pranav Narayanan, et al.
Veröffentlicht: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
von: Bai, Yuqi, et al.
Veröffentlicht: (2025)
von: Bai, Yuqi, et al.
Veröffentlicht: (2025)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
von: Derner, Erik, et al.
Veröffentlicht: (2026)
von: Derner, Erik, et al.
Veröffentlicht: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
von: Bojic, Ljubisa, et al.
Veröffentlicht: (2026)
von: Bojic, Ljubisa, et al.
Veröffentlicht: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
AuditWen:An Open-Source Large Language Model for Audit
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
von: Greco, Candida M., et al.
Veröffentlicht: (2026)
von: Greco, Candida M., et al.
Veröffentlicht: (2026)
Patterns in the Transition From Founder-Leadership to Community Governance of Open Source
von: Noori, Mobina, et al.
Veröffentlicht: (2025)
von: Noori, Mobina, et al.
Veröffentlicht: (2025)
Synthetic Reader Panels: Tournament-Based Ideation with LLM Personas for Autonomous Publishing
von: Zimmerman, Fred
Veröffentlicht: (2026)
von: Zimmerman, Fred
Veröffentlicht: (2026)
LLM Generated Persona is a Promise with a Catch
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
von: Chen, Jiaqi, et al.
Veröffentlicht: (2026)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2026)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
von: Koutcheme, Charles, et al.
Veröffentlicht: (2024)
von: Koutcheme, Charles, et al.
Veröffentlicht: (2024)
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
von: Wu, Yongchao, et al.
Veröffentlicht: (2026)
von: Wu, Yongchao, et al.
Veröffentlicht: (2026)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
von: Fu, Jiachen, et al.
Veröffentlicht: (2025)
von: Fu, Jiachen, et al.
Veröffentlicht: (2025)
Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation
von: Kang, Taewoo, et al.
Veröffentlicht: (2025)
von: Kang, Taewoo, et al.
Veröffentlicht: (2025)
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
von: Zhang, Jingshen, et al.
Veröffentlicht: (2024)
von: Zhang, Jingshen, et al.
Veröffentlicht: (2024)
BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
von: Pai, Tsung-Min, et al.
Veröffentlicht: (2025)
von: Pai, Tsung-Min, et al.
Veröffentlicht: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2024)
Toward Preference-aligned Large Language Models via Residual-based Model Steering
von: La Cava, Lucio, et al.
Veröffentlicht: (2025)
von: La Cava, Lucio, et al.
Veröffentlicht: (2025)
AI-Mediated Communication Can Steer Collective Opinion
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2026)
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
von: Momcilovic, Tomas Bueno, et al.
Veröffentlicht: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025) -
Mechanism of Task-oriented Information Removal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025) -
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025) -
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
von: Costa, Davi Bastos, et al.
Veröffentlicht: (2025) -
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
von: Rupprecht, Jens, et al.
Veröffentlicht: (2025)