Steering Code LLMs with Activation Directions for Language and Library Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rahman, Md Mahbubur, Guha, Arjun, Menon, Harshitha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
von: Zi, Yangtian, et al.
Veröffentlicht: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
Emotion Detection From Social Media Posts
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
Activation Steering with a Feedback Controller
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs
von: Cassano, Federico, et al.
Veröffentlicht: (2023)
von: Cassano, Federico, et al.
Veröffentlicht: (2023)
Steering Language Models With Activation Engineering
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
von: Turner, Alexander Matt, et al.
Veröffentlicht: (2023)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
Depth-Wise Activation Steering for Honest Language Models
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
von: Góral, Gracjan, et al.
Veröffentlicht: (2025)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Elevating Intrusion Detection and Security Fortification in Intelligent Networks through Cutting-Edge Machine Learning Paradigms
von: Munna, Md Minhazul Islam, et al.
Veröffentlicht: (2025)
von: Munna, Md Minhazul Islam, et al.
Veröffentlicht: (2025)
Automated Neuron Labelling Enables Generative Steering and Interpretability in Protein Language Models
von: Banerjee, Arjun, et al.
Veröffentlicht: (2025)
von: Banerjee, Arjun, et al.
Veröffentlicht: (2025)
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
von: Jiang, Xinyan, et al.
Veröffentlicht: (2026)
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
von: Kang, Minjae, et al.
Veröffentlicht: (2026)
von: Kang, Minjae, et al.
Veröffentlicht: (2026)
Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
von: Boruch-Gruszecki, Aleksander, et al.
Veröffentlicht: (2025)
von: Boruch-Gruszecki, Aleksander, et al.
Veröffentlicht: (2025)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Endogenous Resistance to Activation Steering in Language Models
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
von: McKenzie, Alex, et al.
Veröffentlicht: (2026)
Angular Steering: Behavior Control via Rotation in Activation Space
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
von: Vu, Hieu M., et al.
Veröffentlicht: (2025)
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
von: Ghosh, Shaona, et al.
Veröffentlicht: (2025)
Spherical Steering: Geometry-Aware Activation Rotation for Language Models
von: You, Zejia, et al.
Veröffentlicht: (2026)
von: You, Zejia, et al.
Veröffentlicht: (2026)
Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control
von: Skifstad, Julian, et al.
Veröffentlicht: (2026)
von: Skifstad, Julian, et al.
Veröffentlicht: (2026)
SteerConf: Steering LLMs for Confidence Elicitation
von: Zhou, Ziang, et al.
Veröffentlicht: (2025)
von: Zhou, Ziang, et al.
Veröffentlicht: (2025)
Dynamically Scaled Activation Steering
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
von: Ferrando, Alex, et al.
Veröffentlicht: (2025)
HyperSteer: Activation Steering at Scale with Hypernetworks
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
von: Sun, Jiuding, et al.
Veröffentlicht: (2025)
A Unifying Human-Centered AI Fairness Framework
von: Rahman, Munshi Mahbubur, et al.
Veröffentlicht: (2025)
von: Rahman, Munshi Mahbubur, et al.
Veröffentlicht: (2025)
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2026)
Steer Like the LLM: Activation Steering that Mimics Prompting
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
von: Heyman, Geert, et al.
Veröffentlicht: (2026)
Minimizing Collateral Damage in Activation Steering
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
Steered LLM Activations are Non-Surjective
von: Mishra, Aayush, et al.
Veröffentlicht: (2026)
von: Mishra, Aayush, et al.
Veröffentlicht: (2026)
Activation Steering for Chain-of-Thought Compression
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2025)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2025)
Improving Instruction-Following in Language Models through Activation Steering
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
Aligning CodeLLMs with Direct Preference Optimization
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
von: Miao, Yibo, et al.
Veröffentlicht: (2024)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
Towards Causal Deep Learning for Vulnerability Detection
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
von: Haider, Zafaryab, et al.
Veröffentlicht: (2025)
von: Haider, Zafaryab, et al.
Veröffentlicht: (2025)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
von: Jin, Zehao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
von: Zi, Yangtian, et al.
Veröffentlicht: (2025) -
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024) -
Emotion Detection From Social Media Posts
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023) -
Substance Beats Style: Why Beginning Students Fail to Code with LLMs
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024) -
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)