On the Limitations of Steering in Language Model Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Niranjan, Chebrolu, Jaidka, Kokil, Yeo, Gerard Christopher |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
par: Chebrolu, Niranjan, et autres
Publié: (2025)
par: Chebrolu, Niranjan, et autres
Publié: (2025)
Conversations: Love Them, Hate Them, Steer Them
par: Chebrolu, Niranjan, et autres
Publié: (2025)
par: Chebrolu, Niranjan, et autres
Publié: (2025)
Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models
par: Yeo, Gerard, et autres
Publié: (2025)
par: Yeo, Gerard, et autres
Publié: (2025)
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models
par: Yeo, Gerard Christopher, et autres
Publié: (2025)
par: Yeo, Gerard Christopher, et autres
Publié: (2025)
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
par: Churina, Svetlana, et autres
Publié: (2024)
par: Churina, Svetlana, et autres
Publié: (2024)
Beyond Text: Leveraging Multi-Task Learning and Cognitive Appraisal Theory for Post-Purchase Intention Analysis
par: Yeo, Gerard Christopher, et autres
Publié: (2024)
par: Yeo, Gerard Christopher, et autres
Publié: (2024)
Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations
par: Furniturewala, Shaz, et autres
Publié: (2026)
par: Furniturewala, Shaz, et autres
Publié: (2026)
Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
par: Churina, Svetlana, et autres
Publié: (2025)
par: Churina, Svetlana, et autres
Publié: (2025)
Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
par: Singh, Sahajpreet, et autres
Publié: (2025)
par: Singh, Sahajpreet, et autres
Publié: (2025)
Impact of Decoding Methods on Human Alignment of Conversational LLMs
par: Furniturewala, Shaz, et autres
Publié: (2024)
par: Furniturewala, Shaz, et autres
Publié: (2024)
Turn-Level Empathy Prediction Using Psychological Indicators
par: Furniturewala, Shaz, et autres
Publié: (2024)
par: Furniturewala, Shaz, et autres
Publié: (2024)
The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias
par: Verma, Preetika, et autres
Publié: (2024)
par: Verma, Preetika, et autres
Publié: (2024)
CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models
par: Nguyen, Son The, et autres
Publié: (2024)
par: Nguyen, Son The, et autres
Publié: (2024)
Reading Between the Lines: How Electronic Nonverbal Cues shape Emotion Decoding
par: Kumar, Taara, et autres
Publié: (2026)
par: Kumar, Taara, et autres
Publié: (2026)
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
par: Wolf, Yotam, et autres
Publié: (2024)
par: Wolf, Yotam, et autres
Publié: (2024)
Fundamental Limitations of Alignment in Large Language Models
par: Wolf, Yotam, et autres
Publié: (2023)
par: Wolf, Yotam, et autres
Publié: (2023)
Probing the Limits of Stylistic Alignment in Vision-Language Models
par: Farajidizaji, Asma, et autres
Publié: (2025)
par: Farajidizaji, Asma, et autres
Publié: (2025)
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
par: Banerjee, Somnath, et autres
Publié: (2025)
par: Banerjee, Somnath, et autres
Publié: (2025)
"Reasoning" with Rhetoric: On the Style-Evidence Tradeoff in LLM-Generated Counter-Arguments
par: Verma, Preetika, et autres
Publié: (2024)
par: Verma, Preetika, et autres
Publié: (2024)
Self-Steering Language Models
par: Grand, Gabriel, et autres
Publié: (2025)
par: Grand, Gabriel, et autres
Publié: (2025)
CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models
par: Li, Mengfan, et autres
Publié: (2026)
par: Li, Mengfan, et autres
Publié: (2026)
Towards Inference-time Category-wise Safety Steering for Large Language Models
par: Bhattacharjee, Amrita, et autres
Publié: (2024)
par: Bhattacharjee, Amrita, et autres
Publié: (2024)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
par: Cheng, Zifeng, et autres
Publié: (2025)
par: Cheng, Zifeng, et autres
Publié: (2025)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
par: Ouyang, Rongxin, et autres
Publié: (2024)
par: Ouyang, Rongxin, et autres
Publié: (2024)
PHAnToM: Persona-based Prompting Has An Effect on Theory-of-Mind Reasoning in Large Language Models
par: Tan, Fiona Anting, et autres
Publié: (2024)
par: Tan, Fiona Anting, et autres
Publié: (2024)
Inference-time Alignment via Sparse Junction Steering
par: Hu, Runyi, et autres
Publié: (2026)
par: Hu, Runyi, et autres
Publié: (2026)
Probing and Steering Evaluation Awareness of Language Models
par: Nguyen, Jord, et autres
Publié: (2025)
par: Nguyen, Jord, et autres
Publié: (2025)
Activation Scaling for Steering and Interpreting Language Models
par: Stoehr, Niklas, et autres
Publié: (2024)
par: Stoehr, Niklas, et autres
Publié: (2024)
PlotTwist: A Creative Plot Generation Framework with Small Language Models
par: Thorat, Abhinav, et autres
Publié: (2026)
par: Thorat, Abhinav, et autres
Publié: (2026)
GitSearch: Enhancing Community Notes Generation with Gap-Informed Targeted Search
par: Singh, Sahajpreet, et autres
Publié: (2026)
par: Singh, Sahajpreet, et autres
Publié: (2026)
Continual Learning with Global Alignment
par: Bai, Xueying, et autres
Publié: (2022)
par: Bai, Xueying, et autres
Publié: (2022)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
par: Braun, Joschka
Publié: (2026)
par: Braun, Joschka
Publié: (2026)
CogSteer: Cognition-Inspired Selective Layer Intervention for Efficiently Steering Large Language Models
par: Wang, Xintong, et autres
Publié: (2024)
par: Wang, Xintong, et autres
Publié: (2024)
Compositional Steering of Large Language Models with Steering Tokens
par: Radevski, Gorjan, et autres
Publié: (2026)
par: Radevski, Gorjan, et autres
Publié: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
Prompt-Based Value Steering of Large Language Models
par: Abbo, Giulio Antonio, et autres
Publié: (2025)
par: Abbo, Giulio Antonio, et autres
Publié: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
par: Pokharel, Rhitabrat, et autres
Publié: (2026)
par: Pokharel, Rhitabrat, et autres
Publié: (2026)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
From What to Why: Thought-Space Recommendation with Small Language Models
par: Biswas, Prosenjit, et autres
Publié: (2025)
par: Biswas, Prosenjit, et autres
Publié: (2025)
Natural Language Interaction with Databases on Edge Devices in the Internet of Battlefield Things
par: Molek, Christopher D., et autres
Publié: (2025)
par: Molek, Christopher D., et autres
Publié: (2025)
Documents similaires
-
From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
par: Chebrolu, Niranjan, et autres
Publié: (2025) -
Conversations: Love Them, Hate Them, Steer Them
par: Chebrolu, Niranjan, et autres
Publié: (2025) -
Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models
par: Yeo, Gerard, et autres
Publié: (2025) -
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models
par: Yeo, Gerard Christopher, et autres
Publié: (2025) -
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
par: Churina, Svetlana, et autres
Publié: (2024)