A Unified Understanding and Evaluation of Steering Methods
Fuente:
arXiv
Salvato in:
| Autori principali: | Im, Shawn, Li, Sharon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026)
di: Im, Shawn, et al.
Pubblicazione: (2026)
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
di: Oh, Changdae, et al.
Pubblicazione: (2025)
di: Oh, Changdae, et al.
Pubblicazione: (2025)
Towards Understanding Steering Strength
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
di: Im, Shawn, et al.
Pubblicazione: (2024)
di: Im, Shawn, et al.
Pubblicazione: (2024)
How Well Can Preference Optimization Generalize Under Noisy Feedback?
di: Im, Shawn, et al.
Pubblicazione: (2025)
di: Im, Shawn, et al.
Pubblicazione: (2025)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
di: Weng, Zixuan, et al.
Pubblicazione: (2026)
di: Weng, Zixuan, et al.
Pubblicazione: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
di: Zhou, Ziang, et al.
Pubblicazione: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
di: Lucchetti, Francesca, et al.
Pubblicazione: (2024)
Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Understanding and Mitigating Dataset Corruption in LLM Steering
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
di: Anderson, Cullen, et al.
Pubblicazione: (2026)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning
di: Li, Wendi, et al.
Pubblicazione: (2026)
di: Li, Wendi, et al.
Pubblicazione: (2026)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2025)
Conceptors for Semantic Steering
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Ilias, et al.
Pubblicazione: (2026)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
Understanding the Learning Dynamics of Alignment with Human Feedback
di: Im, Shawn, et al.
Pubblicazione: (2024)
di: Im, Shawn, et al.
Pubblicazione: (2024)
That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
di: Bae, Jaesung, et al.
Pubblicazione: (2025)
di: Bae, Jaesung, et al.
Pubblicazione: (2025)
A Unified Benchmark for Evaluating Knowledge Graph Construction Methods and Graph Neural Networks
di: Kabal, Othmane, et al.
Pubblicazione: (2026)
di: Kabal, Othmane, et al.
Pubblicazione: (2026)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
di: Cao, Yuanpu, et al.
Pubblicazione: (2024)
A Unified Study of LoRA Variants: Taxonomy, Review, Codebase, and Empirical Evaluation
di: He, Haonan, et al.
Pubblicazione: (2026)
di: He, Haonan, et al.
Pubblicazione: (2026)
ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
di: Avrahami, Eden, et al.
Pubblicazione: (2026)
di: Avrahami, Eden, et al.
Pubblicazione: (2026)
Steering Language Models with Weight Arithmetic
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
di: Fierro, Constanza, et al.
Pubblicazione: (2025)
Predicting Where Steering Vectors Succeed
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Steering Language Models With Activation Engineering
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
di: Turner, Alexander Matt, et al.
Pubblicazione: (2023)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
di: Braun, Joschka
Pubblicazione: (2026)
di: Braun, Joschka
Pubblicazione: (2026)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
di: Li, Zihao, et al.
Pubblicazione: (2025)
di: Li, Zihao, et al.
Pubblicazione: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
di: Heyman, Geert, et al.
Pubblicazione: (2026)
di: Heyman, Geert, et al.
Pubblicazione: (2026)
Compositional Steering of Large Language Models with Steering Tokens
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
di: Radevski, Gorjan, et al.
Pubblicazione: (2026)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
di: Cheng, Stephen, et al.
Pubblicazione: (2026)
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Evaluating Copyright Takedown Methods for Language Models
di: Wei, Boyi, et al.
Pubblicazione: (2024)
di: Wei, Boyi, et al.
Pubblicazione: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
di: Gan, Woody Haosheng, et al.
Pubblicazione: (2025)
Differentially Private Steering for Large Language Model Alignment
di: Goel, Anmol, et al.
Pubblicazione: (2025)
di: Goel, Anmol, et al.
Pubblicazione: (2025)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
di: Zhang, Xuechen, et al.
Pubblicazione: (2026)
ROAST: Rollout-based On-distribution Activation Steering Technique
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
di: Su, Xuanbo, et al.
Pubblicazione: (2026)
Social Caption: Evaluating Social Understanding in Multimodal Models
di: Thumu, Bhaavanaa, et al.
Pubblicazione: (2026)
di: Thumu, Bhaavanaa, et al.
Pubblicazione: (2026)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
di: Im, Shawn, et al.
Pubblicazione: (2026) -
Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
di: Oh, Changdae, et al.
Pubblicazione: (2025) -
Towards Understanding Steering Strength
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026) -
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
di: Im, Shawn, et al.
Pubblicazione: (2024) -
How Well Can Preference Optimization Generalize Under Noisy Feedback?
di: Im, Shawn, et al.
Pubblicazione: (2025)