ROAST: Rollout-based On-distribution Activation Steering Technique
Fuente:
arXiv
Guardado en:
| Autores principales: | Su, Xuanbo, Luo, Hao, Zhang, Yingfang, Zhang, Lijun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
por: Jin, Zehao, et al.
Publicado: (2026)
por: Jin, Zehao, et al.
Publicado: (2026)
Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
por: Su, Xuanbo, et al.
Publicado: (2025)
por: Su, Xuanbo, et al.
Publicado: (2025)
ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection for ABSA
por: Chebolu, Siva Uday Sampreeth, et al.
Publicado: (2024)
por: Chebolu, Siva Uday Sampreeth, et al.
Publicado: (2024)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
por: Luo, Sijia, et al.
Publicado: (2026)
por: Luo, Sijia, et al.
Publicado: (2026)
Steering Language Models With Activation Engineering
por: Turner, Alexander Matt, et al.
Publicado: (2023)
por: Turner, Alexander Matt, et al.
Publicado: (2023)
HyperSteer: Activation Steering at Scale with Hypernetworks
por: Sun, Jiuding, et al.
Publicado: (2025)
por: Sun, Jiuding, et al.
Publicado: (2025)
Steer Like the LLM: Activation Steering that Mimics Prompting
por: Heyman, Geert, et al.
Publicado: (2026)
por: Heyman, Geert, et al.
Publicado: (2026)
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
por: Yue, Yanwei, et al.
Publicado: (2026)
por: Yue, Yanwei, et al.
Publicado: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
por: Liu, Bingshuai, et al.
Publicado: (2025)
por: Liu, Bingshuai, et al.
Publicado: (2025)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
por: You, Zejia, et al.
Publicado: (2026)
por: You, Zejia, et al.
Publicado: (2026)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
por: Hegazy, Amr, et al.
Publicado: (2025)
por: Hegazy, Amr, et al.
Publicado: (2025)
Steering MoE LLMs via Expert (De)Activation
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
por: Fayyaz, Mohsen, et al.
Publicado: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025)
por: Xu, Yixuan Even, et al.
Publicado: (2025)
Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling
por: Wang, Xinglin, et al.
Publicado: (2026)
por: Wang, Xinglin, et al.
Publicado: (2026)
Programming Refusal with Conditional Activation Steering
por: Lee, Bruce W., et al.
Publicado: (2024)
por: Lee, Bruce W., et al.
Publicado: (2024)
SAKE: Steering Activations for Knowledge Editing
por: Scialanga, Marco, et al.
Publicado: (2025)
por: Scialanga, Marco, et al.
Publicado: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
por: Lucchetti, Francesca, et al.
Publicado: (2024)
por: Lucchetti, Francesca, et al.
Publicado: (2024)
FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
por: Wang, Sihan, et al.
Publicado: (2026)
por: Wang, Sihan, et al.
Publicado: (2026)
HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
por: Wang, Xinyi, et al.
Publicado: (2025)
por: Wang, Xinyi, et al.
Publicado: (2025)
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
por: Bao, Yuntai, et al.
Publicado: (2026)
por: Bao, Yuntai, et al.
Publicado: (2026)
Endogenous Resistance to Activation Steering in Language Models
por: McKenzie, Alex, et al.
Publicado: (2026)
por: McKenzie, Alex, et al.
Publicado: (2026)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
por: Deshpande, Vijeta, et al.
Publicado: (2026)
por: Deshpande, Vijeta, et al.
Publicado: (2026)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
por: Zhang, Yi-Kai, et al.
Publicado: (2026)
por: Zhang, Yi-Kai, et al.
Publicado: (2026)
Extracting Unlearned Information from LLMs with Activation Steering
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
por: Seyitoğlu, Atakan, et al.
Publicado: (2024)
Steering Llama 2 via Contrastive Activation Addition
por: Panickssery, Nina, et al.
Publicado: (2023)
por: Panickssery, Nina, et al.
Publicado: (2023)
Activation Steering via Generative Causal Mediation
por: Sankaranarayanan, Aruna, et al.
Publicado: (2026)
por: Sankaranarayanan, Aruna, et al.
Publicado: (2026)
VSPO: Vector-Steered Policy Optimization for Behavioral Control
por: Zhang, Xuechen, et al.
Publicado: (2026)
por: Zhang, Xuechen, et al.
Publicado: (2026)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
por: Ding, Yi, et al.
Publicado: (2026)
por: Ding, Yi, et al.
Publicado: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
por: van der Weij, Teun, et al.
Publicado: (2024)
por: van der Weij, Teun, et al.
Publicado: (2024)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
por: Cao, Yuanpu, et al.
Publicado: (2024)
por: Cao, Yuanpu, et al.
Publicado: (2024)
Improving Instruction-Following in Language Models through Activation Steering
por: Stolfo, Alessandro, et al.
Publicado: (2024)
por: Stolfo, Alessandro, et al.
Publicado: (2024)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
por: Iso, Hayate, et al.
Publicado: (2026)
por: Iso, Hayate, et al.
Publicado: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
por: Zhou, Ziang, et al.
Publicado: (2025)
por: Zhou, Ziang, et al.
Publicado: (2025)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
por: Anand, Nikhil, et al.
Publicado: (2026)
por: Anand, Nikhil, et al.
Publicado: (2026)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
por: Scalena, Daniel, et al.
Publicado: (2024)
por: Scalena, Daniel, et al.
Publicado: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
por: Bigelow, Eric, et al.
Publicado: (2025)
por: Bigelow, Eric, et al.
Publicado: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
por: Soo, Samuel, et al.
Publicado: (2025)
por: Soo, Samuel, et al.
Publicado: (2025)
Conceptors for Semantic Steering
por: Triantafyllopoulos, Ilias, et al.
Publicado: (2026)
por: Triantafyllopoulos, Ilias, et al.
Publicado: (2026)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
por: Xi, Haocheng, et al.
Publicado: (2026)
por: Xi, Haocheng, et al.
Publicado: (2026)
Ejemplares similares
-
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
por: Jin, Zehao, et al.
Publicado: (2026) -
Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
por: Su, Xuanbo, et al.
Publicado: (2025) -
ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection for ABSA
por: Chebolu, Siva Uday Sampreeth, et al.
Publicado: (2024) -
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
por: Luo, Sijia, et al.
Publicado: (2026) -
Steering Language Models With Activation Engineering
por: Turner, Alexander Matt, et al.
Publicado: (2023)