Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
Fuente:
arXiv
Guardado en:
| Autores principales: | Kang, Minjae, Kim, Jaehyung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Instruction-Following in Language Models through Activation Steering
por: Stolfo, Alessandro, et al.
Publicado: (2024)
por: Stolfo, Alessandro, et al.
Publicado: (2024)
Few-shot Personalization of LLMs with Mis-aligned Responses
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Dynamically Scaled Activation Steering
por: Ferrando, Alex, et al.
Publicado: (2025)
por: Ferrando, Alex, et al.
Publicado: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
por: Nam, Jaehyun, et al.
Publicado: (2024)
por: Nam, Jaehyun, et al.
Publicado: (2024)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
por: Seo, Yeongbin, et al.
Publicado: (2025)
por: Seo, Yeongbin, et al.
Publicado: (2025)
Debiasing Online Preference Learning via Preference Feature Preservation
por: Kim, Dongyoung, et al.
Publicado: (2025)
por: Kim, Dongyoung, et al.
Publicado: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
por: Lee, Dongseok, et al.
Publicado: (2025)
por: Lee, Dongseok, et al.
Publicado: (2025)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
por: Jiang, Xinyan, et al.
Publicado: (2026)
por: Jiang, Xinyan, et al.
Publicado: (2026)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
por: Ismail, Ahmed H., et al.
Publicado: (2026)
por: Ismail, Ahmed H., et al.
Publicado: (2026)
Extracting Unlearned Information from LLMs with Activation Steering
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
Steering LLMs via Scalable Interactive Oversight
por: Zhou, Enyu, et al.
Publicado: (2026)
por: Zhou, Enyu, et al.
Publicado: (2026)
Self-Evolving LLMs via Continual Instruction Tuning
por: Kang, Jiazheng, et al.
Publicado: (2025)
por: Kang, Jiazheng, et al.
Publicado: (2025)
Angular Steering: Behavior Control via Rotation in Activation Space
por: Vu, Hieu M., et al.
Publicado: (2025)
por: Vu, Hieu M., et al.
Publicado: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
por: He, Zirui, et al.
Publicado: (2025)
por: He, Zirui, et al.
Publicado: (2025)
Structural Reasoning Improves Molecular Understanding of LLM
por: Jang, Yunhui, et al.
Publicado: (2024)
por: Jang, Yunhui, et al.
Publicado: (2024)
Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
por: Nam, Yunhun, et al.
Publicado: (2025)
por: Nam, Yunhun, et al.
Publicado: (2025)
Minimizing Collateral Damage in Activation Steering
por: Nguyen, Tam, et al.
Publicado: (2026)
por: Nguyen, Tam, et al.
Publicado: (2026)
Steered LLM Activations are Non-Surjective
por: Mishra, Aayush, et al.
Publicado: (2026)
por: Mishra, Aayush, et al.
Publicado: (2026)
Activation Steering for Chain-of-Thought Compression
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
por: Azizi, Seyedarmin, et al.
Publicado: (2025)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
por: Kwak, Minseo, et al.
Publicado: (2026)
por: Kwak, Minseo, et al.
Publicado: (2026)
Hierarchical Meta-Reinforcement Learning via Automated Macro-Action Discovery
por: Cho, Minjae, et al.
Publicado: (2024)
por: Cho, Minjae, et al.
Publicado: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
por: Sun, Jiuding, et al.
Publicado: (2025)
por: Sun, Jiuding, et al.
Publicado: (2025)
Enhancing and Assessing Instruction-Following with Fine-Grained Instruction Variants
por: Yang, Jiuding, et al.
Publicado: (2024)
por: Yang, Jiuding, et al.
Publicado: (2024)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
por: Han, Pengrui, et al.
Publicado: (2026)
por: Han, Pengrui, et al.
Publicado: (2026)
The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
por: Kang, Shijia, et al.
Publicado: (2025)
por: Kang, Shijia, et al.
Publicado: (2025)
Steering Llama 2 via Contrastive Activation Addition
por: Panickssery, Nina, et al.
Publicado: (2023)
por: Panickssery, Nina, et al.
Publicado: (2023)
Steer Like the LLM: Activation Steering that Mimics Prompting
por: Heyman, Geert, et al.
Publicado: (2026)
por: Heyman, Geert, et al.
Publicado: (2026)
Tabular Transfer Learning via Prompting LLMs
por: Nam, Jaehyun, et al.
Publicado: (2024)
por: Nam, Jaehyun, et al.
Publicado: (2024)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
por: Song, Woomin, et al.
Publicado: (2024)
por: Song, Woomin, et al.
Publicado: (2024)
Online Continual Learning For Interactive Instruction Following Agents
por: Kim, Byeonghwi, et al.
Publicado: (2024)
por: Kim, Byeonghwi, et al.
Publicado: (2024)
Personalized LLM Decoding via Contrasting Personal Preference
por: Bu, Hyungjune, et al.
Publicado: (2025)
por: Bu, Hyungjune, et al.
Publicado: (2025)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
por: Korznikov, Anton, et al.
Publicado: (2025)
por: Korznikov, Anton, et al.
Publicado: (2025)
Test-time Diverse Reasoning by Riemannian Activation Steering
por: Khanh, Ly Tran Ho, et al.
Publicado: (2025)
por: Khanh, Ly Tran Ho, et al.
Publicado: (2025)
Steering Large Language Model Activations in Sparse Spaces
por: Bayat, Reza, et al.
Publicado: (2025)
por: Bayat, Reza, et al.
Publicado: (2025)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
por: Wang, Chenyang, et al.
Publicado: (2025)
por: Wang, Chenyang, et al.
Publicado: (2025)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
por: Zhang, Zeyu, et al.
Publicado: (2026)
por: Zhang, Zeyu, et al.
Publicado: (2026)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
por: Scalena, Daniel, et al.
Publicado: (2024)
por: Scalena, Daniel, et al.
Publicado: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
por: Bigelow, Eric, et al.
Publicado: (2025)
por: Bigelow, Eric, et al.
Publicado: (2025)
Ejemplares similares
-
Improving Instruction-Following in Language Models through Activation Steering
por: Stolfo, Alessandro, et al.
Publicado: (2024) -
Few-shot Personalization of LLMs with Mis-aligned Responses
por: Kim, Jaehyung, et al.
Publicado: (2024) -
Dynamically Scaled Activation Steering
por: Ferrando, Alex, et al.
Publicado: (2025) -
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024) -
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
por: Nam, Jaehyun, et al.
Publicado: (2024)