Steering at the Source: Style Modulation Heads for Robust Persona Control
Fuente:
arXiv
Guardado en:
| Autores principales: | Izawa, Yoshihiro, Minegishi, Gouki, Eguchi, Koshi, Hosokawa, Sosuke, Taura, Kenjiro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
por: Cho, Hakaze, et al.
Publicado: (2025)
por: Cho, Hakaze, et al.
Publicado: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025)
por: Minegishi, Gouki, et al.
Publicado: (2025)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
por: Costa, Davi Bastos, et al.
Publicado: (2025)
por: Costa, Davi Bastos, et al.
Publicado: (2025)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
por: Rupprecht, Jens, et al.
Publicado: (2025)
por: Rupprecht, Jens, et al.
Publicado: (2025)
Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
por: Gelberg, Yoav, et al.
Publicado: (2025)
por: Gelberg, Yoav, et al.
Publicado: (2025)
The Need for a Socially-Grounded Persona Framework for User Simulation
por: Venkit, Pranav Narayanan, et al.
Publicado: (2026)
por: Venkit, Pranav Narayanan, et al.
Publicado: (2026)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
por: Kim, Jiseon, et al.
Publicado: (2025)
por: Kim, Jiseon, et al.
Publicado: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
por: Kiet, Huynh Trung, et al.
Publicado: (2026)
por: Kiet, Huynh Trung, et al.
Publicado: (2026)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
por: Venkit, Pranav Narayanan, et al.
Publicado: (2025)
por: Venkit, Pranav Narayanan, et al.
Publicado: (2025)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
por: Kabir, Mohsinul, et al.
Publicado: (2025)
por: Kabir, Mohsinul, et al.
Publicado: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
por: Bai, Yuqi, et al.
Publicado: (2025)
por: Bai, Yuqi, et al.
Publicado: (2025)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
por: Tzachristas, Ioannis, et al.
Publicado: (2025)
por: Tzachristas, Ioannis, et al.
Publicado: (2025)
Mind the Style: Impact of Communication Style on Human-Chatbot Interaction
por: Derner, Erik, et al.
Publicado: (2026)
por: Derner, Erik, et al.
Publicado: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
por: Xiao, Yuxin, et al.
Publicado: (2025)
por: Xiao, Yuxin, et al.
Publicado: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
por: Bojic, Ljubisa, et al.
Publicado: (2026)
por: Bojic, Ljubisa, et al.
Publicado: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
por: van der Weij, Teun, et al.
Publicado: (2024)
por: van der Weij, Teun, et al.
Publicado: (2024)
AuditWen:An Open-Source Large Language Model for Audit
por: Huang, Jiajia, et al.
Publicado: (2024)
por: Huang, Jiajia, et al.
Publicado: (2024)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
por: Greco, Candida M., et al.
Publicado: (2026)
por: Greco, Candida M., et al.
Publicado: (2026)
Patterns in the Transition From Founder-Leadership to Community Governance of Open Source
por: Noori, Mobina, et al.
Publicado: (2025)
por: Noori, Mobina, et al.
Publicado: (2025)
Synthetic Reader Panels: Tournament-Based Ideation with LLM Personas for Autonomous Publishing
por: Zimmerman, Fred
Publicado: (2026)
por: Zimmerman, Fred
Publicado: (2026)
LLM Generated Persona is a Promise with a Catch
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
por: Chen, Jiaqi, et al.
Publicado: (2026)
por: Chen, Jiaqi, et al.
Publicado: (2026)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
por: Koutcheme, Charles, et al.
Publicado: (2024)
por: Koutcheme, Charles, et al.
Publicado: (2024)
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
por: Wu, Yongchao, et al.
Publicado: (2026)
por: Wu, Yongchao, et al.
Publicado: (2026)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
por: Walsh, Cole, et al.
Publicado: (2026)
por: Walsh, Cole, et al.
Publicado: (2026)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
por: Fu, Jiachen, et al.
Publicado: (2025)
por: Fu, Jiachen, et al.
Publicado: (2025)
Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation
por: Kang, Taewoo, et al.
Publicado: (2025)
por: Kang, Taewoo, et al.
Publicado: (2025)
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
por: Sun, Lihao, et al.
Publicado: (2026)
por: Sun, Lihao, et al.
Publicado: (2026)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
por: Truong, Kimberly Le, et al.
Publicado: (2025)
por: Truong, Kimberly Le, et al.
Publicado: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
por: Zhang, Jingshen, et al.
Publicado: (2024)
por: Zhang, Jingshen, et al.
Publicado: (2024)
BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
por: Pai, Tsung-Min, et al.
Publicado: (2025)
por: Pai, Tsung-Min, et al.
Publicado: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
por: Burdisso, Sergio, et al.
Publicado: (2024)
por: Burdisso, Sergio, et al.
Publicado: (2024)
Toward Preference-aligned Large Language Models via Residual-based Model Steering
por: La Cava, Lucio, et al.
Publicado: (2025)
por: La Cava, Lucio, et al.
Publicado: (2025)
AI-Mediated Communication Can Steer Collective Opinion
por: Tsirtsis, Stratis, et al.
Publicado: (2026)
por: Tsirtsis, Stratis, et al.
Publicado: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
por: Ngueajio, Mikel K., et al.
Publicado: (2025)
por: Ngueajio, Mikel K., et al.
Publicado: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
por: Zheng, Mingqian, et al.
Publicado: (2023)
por: Zheng, Mingqian, et al.
Publicado: (2023)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
por: Momcilovic, Tomas Bueno, et al.
Publicado: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
por: Joshi, Abhinav, et al.
Publicado: (2024)
por: Joshi, Abhinav, et al.
Publicado: (2024)
Ejemplares similares
-
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
por: Minegishi, Gouki, et al.
Publicado: (2025) -
Mechanism of Task-oriented Information Removal in In-context Learning
por: Cho, Hakaze, et al.
Publicado: (2025) -
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
por: Minegishi, Gouki, et al.
Publicado: (2025) -
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
por: Costa, Davi Bastos, et al.
Publicado: (2025) -
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
por: Rupprecht, Jens, et al.
Publicado: (2025)