Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Fuente:
arXiv
Saved in:
| Main Authors: | Herbster, Niklas, Zborowski, Martin, Tosato, Alberto, Gidel, Gauthier, Tosato, Tommaso |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
by: Tosato, Tommaso, et al.
Published: (2025)
by: Tosato, Tommaso, et al.
Published: (2025)
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
by: Tosato, Tommaso, et al.
Published: (2024)
by: Tosato, Tommaso, et al.
Published: (2024)
TUM-MiKaNi at SemEval-2025 Task 3: Towards Multilingual and Knowledge-Aware Non-factual Hallucination Identification
by: Anschütz, Miriam, et al.
Published: (2025)
by: Anschütz, Miriam, et al.
Published: (2025)
Subword Embedding from Bytes Gains Privacy without Sacrificing Accuracy and Complexity
by: Zhang, Mengjiao, et al.
Published: (2024)
by: Zhang, Mengjiao, et al.
Published: (2024)
CONHECIMENTO DO INDIVÍDUO OSTOMIZADO EM RELAÇÃO AO AUTOCUIDADO
by: Sonia Ramos Tosato
Published: (2006)
by: Sonia Ramos Tosato
Published: (2006)
Accelerating battery research with an AI interface between FINALES and Kadi4Mat
by: Tosato, Giovanna, et al.
Published: (2026)
by: Tosato, Giovanna, et al.
Published: (2026)
Activation Scaling for Steering and Interpreting Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences
by: Ferbach, Damien, et al.
Published: (2024)
by: Ferbach, Damien, et al.
Published: (2024)
Tight Lower Bounds and Improved Convergence in Performative Prediction
by: Khorsandi, Pedram, et al.
Published: (2024)
by: Khorsandi, Pedram, et al.
Published: (2024)
GraphTranslator: Aligning Graph Model to Large Language Model for Open-ended Tasks
by: Zhang, Mengmei, et al.
Published: (2024)
by: Zhang, Mengmei, et al.
Published: (2024)
Learning Page Order in Shuffled WOO Releases
by: Kahraman, Efe, et al.
Published: (2026)
by: Kahraman, Efe, et al.
Published: (2026)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
by: Xia, Wei, et al.
Published: (2025)
by: Xia, Wei, et al.
Published: (2025)
FaithSteer-BENCH: A Deployment-Aligned Stress-Testing Benchmark for Inference-Time Steering
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
Dynamically Scaled Activation Steering
by: Ferrando, Alex, et al.
Published: (2025)
by: Ferrando, Alex, et al.
Published: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by: Waghjale, Siddhant, et al.
Published: (2024)
by: Waghjale, Siddhant, et al.
Published: (2024)
Genre Controlled Music Generation via Activation Steering
by: Narashiman, Swathi, et al.
Published: (2025)
by: Narashiman, Swathi, et al.
Published: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
by: Heyman, Geert, et al.
Published: (2026)
by: Heyman, Geert, et al.
Published: (2026)
Minimizing Collateral Damage in Activation Steering
by: Nguyen, Tam, et al.
Published: (2026)
by: Nguyen, Tam, et al.
Published: (2026)
Steered LLM Activations are Non-Surjective
by: Mishra, Aayush, et al.
Published: (2026)
by: Mishra, Aayush, et al.
Published: (2026)
Activation Steering for Chain-of-Thought Compression
by: Azizi, Seyedarmin, et al.
Published: (2025)
by: Azizi, Seyedarmin, et al.
Published: (2025)
Endogenous Resistance to Activation Steering in Language Models
by: McKenzie, Alex, et al.
Published: (2026)
by: McKenzie, Alex, et al.
Published: (2026)
MidSteer: Optimal Affine Framework for Steering Generative Models
by: Gaintseva, Tatiana, et al.
Published: (2026)
by: Gaintseva, Tatiana, et al.
Published: (2026)
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering
by: Ye, Wencheng, et al.
Published: (2026)
by: Ye, Wencheng, et al.
Published: (2026)
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
by: Dubey, Shivam
Published: (2025)
by: Dubey, Shivam
Published: (2025)
Fusion Steering: Prompt-Specific Activation Control
by: Chang, Waldemar, et al.
Published: (2025)
by: Chang, Waldemar, et al.
Published: (2025)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
by: Jiang, Xinyan, et al.
Published: (2026)
by: Jiang, Xinyan, et al.
Published: (2026)
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
by: Sui, Yuan, et al.
Published: (2026)
by: Sui, Yuan, et al.
Published: (2026)
YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation
by: Bounhar, Abdelaziz, et al.
Published: (2026)
by: Bounhar, Abdelaziz, et al.
Published: (2026)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
by: Kang, Diancheng, et al.
Published: (2026)
by: Kang, Diancheng, et al.
Published: (2026)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
by: Xiong, Chen, et al.
Published: (2026)
by: Xiong, Chen, et al.
Published: (2026)
Programming Refusal with Conditional Activation Steering
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
SAKE: Steering Activations for Knowledge Editing
by: Scialanga, Marco, et al.
Published: (2025)
by: Scialanga, Marco, et al.
Published: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
by: Ismail, Ahmed H., et al.
Published: (2026)
by: Ismail, Ahmed H., et al.
Published: (2026)
Similar Items
-
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History
by: Tosato, Tommaso, et al.
Published: (2025) -
Lost in Translation: The Algorithmic Gap Between LMs and the Brain
by: Tosato, Tommaso, et al.
Published: (2024) -
TUM-MiKaNi at SemEval-2025 Task 3: Towards Multilingual and Knowledge-Aware Non-factual Hallucination Identification
by: Anschütz, Miriam, et al.
Published: (2025) -
Subword Embedding from Bytes Gains Privacy without Sacrificing Accuracy and Complexity
by: Zhang, Mengjiao, et al.
Published: (2024) -
CONHECIMENTO DO INDIVÍDUO OSTOMIZADO EM RELAÇÃO AO AUTOCUIDADO
by: Sonia Ramos Tosato
Published: (2006)