Exploring the Personality Traits of LLMs through Latent Features Steering
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Shu, Zhu, Shenzhe, Liu, Liang, Hu, Lijie, Li, Mengdi, Wang, Di |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
di: Zhu, Shenzhe
Pubblicazione: (2025)
di: Zhu, Shenzhe
Pubblicazione: (2025)
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
di: Yang, Shu, et al.
Pubblicazione: (2024)
di: Yang, Shu, et al.
Pubblicazione: (2024)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Mitigating Misalignment Contagion by Steering with Implicit Traits
di: Chang, Maria, et al.
Pubblicazione: (2026)
di: Chang, Maria, et al.
Pubblicazione: (2026)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
di: Lv, Hang, et al.
Pubblicazione: (2025)
di: Lv, Hang, et al.
Pubblicazione: (2025)
A Hopfieldian View-based Interpretation for Chain-of-Thought Reasoning
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
The Dark Patterns of Personalized Persuasion in Large Language Models: Exposing Persuasive Linguistic Features for Big Five Personality Traits in LLMs Responses
di: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Pubblicazione: (2024)
di: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Pubblicazione: (2024)
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
di: Allbert, Rumi, et al.
Pubblicazione: (2024)
di: Allbert, Rumi, et al.
Pubblicazione: (2024)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
di: Zhou, Wenrui, et al.
Pubblicazione: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
di: Wang, Anyi, et al.
Pubblicazione: (2025)
di: Wang, Anyi, et al.
Pubblicazione: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
di: Zhang, Ruikang, et al.
Pubblicazione: (2026)
PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
di: Yu, Jiongchi, et al.
Pubblicazione: (2026)
di: Yu, Jiongchi, et al.
Pubblicazione: (2026)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
Context Steering: Controllable Personalization at Inference Time
di: He, Jerry Zhi-Yang, et al.
Pubblicazione: (2024)
di: He, Jerry Zhi-Yang, et al.
Pubblicazione: (2024)
A Comparative Study of Large Language Models and Human Personality Traits
di: Jiaqi, Wang, et al.
Pubblicazione: (2025)
di: Jiaqi, Wang, et al.
Pubblicazione: (2025)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
di: You, Liangliang, et al.
Pubblicazione: (2025)
di: You, Liangliang, et al.
Pubblicazione: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
di: Wang, Shu, et al.
Pubblicazione: (2024)
di: Wang, Shu, et al.
Pubblicazione: (2024)
The Compositional Architecture of Regret in Large Language Models
di: Cui, Xiangxiang, et al.
Pubblicazione: (2025)
di: Cui, Xiangxiang, et al.
Pubblicazione: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Ting, et al.
Pubblicazione: (2025)
Contextual Categorization Enhancement through LLMs Latent-Space
di: Bettouche, Zineddine, et al.
Pubblicazione: (2024)
di: Bettouche, Zineddine, et al.
Pubblicazione: (2024)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
Eliciting Personality Traits in Large Language Models
di: Hilliard, Airlie, et al.
Pubblicazione: (2024)
di: Hilliard, Airlie, et al.
Pubblicazione: (2024)
Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework
di: Ma, Xilai, et al.
Pubblicazione: (2026)
di: Ma, Xilai, et al.
Pubblicazione: (2026)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
Predicting the Big Five Personality Traits in Chinese Counselling Dialogues Using Large Language Models
di: Yan, Yang, et al.
Pubblicazione: (2024)
di: Yan, Yang, et al.
Pubblicazione: (2024)
On Effects of Steering Latent Representation for Large Language Model Unlearning
di: Huu-Tien, Dang, et al.
Pubblicazione: (2024)
di: Huu-Tien, Dang, et al.
Pubblicazione: (2024)
Understanding the Repeat Curse in Large Language Models from a Feature Perspective
di: Yao, Junchi, et al.
Pubblicazione: (2025)
di: Yao, Junchi, et al.
Pubblicazione: (2025)
SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control
di: Chittem, Adithya, et al.
Pubblicazione: (2025)
di: Chittem, Adithya, et al.
Pubblicazione: (2025)
HiURE: Hierarchical Exemplar Contrastive Learning for Unsupervised Relation Extraction
di: Liu, Shuliang, et al.
Pubblicazione: (2022)
di: Liu, Shuliang, et al.
Pubblicazione: (2022)
ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
di: Guang, Jiahui, et al.
Pubblicazione: (2026)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
di: Tsaknakis, Ioannis, et al.
Pubblicazione: (2025)
di: Tsaknakis, Ioannis, et al.
Pubblicazione: (2025)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
di: Banayeeanzade, Amin, et al.
Pubblicazione: (2025)
KV Cache Steering for Controlling Frozen LLMs
di: Belitsky, Max, et al.
Pubblicazione: (2025)
di: Belitsky, Max, et al.
Pubblicazione: (2025)
When Only the Final Text Survives: Implicit Execution Tracing for Multi-Agent Attribution
di: Nian, Yi, et al.
Pubblicazione: (2026)
di: Nian, Yi, et al.
Pubblicazione: (2026)
Private Language Models via Truncated Laplacian Mechanism
di: Huang, Tianhao, et al.
Pubblicazione: (2024)
di: Huang, Tianhao, et al.
Pubblicazione: (2024)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
di: Pres, Itamar, et al.
Pubblicazione: (2024)
di: Pres, Itamar, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs
di: Yang, Shu, et al.
Pubblicazione: (2024) -
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
di: Zhu, Shenzhe
Pubblicazione: (2025) -
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning
di: Yang, Shu, et al.
Pubblicazione: (2024) -
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024) -
Mitigating Misalignment Contagion by Steering with Implicit Traits
di: Chang, Maria, et al.
Pubblicazione: (2026)