Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Seonglae, Wu, Zekun, Koshiyama, Adriano |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026)
by: Patel, Het, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
by: Zhang, Luyan
Published: (2025)
by: Zhang, Luyan
Published: (2025)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)
by: Resck, Lucas, et al.
Published: (2026)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
by: Gwak, Jiho, et al.
Published: (2025)
by: Gwak, Jiho, et al.
Published: (2025)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
by: He, Langzhou, et al.
Published: (2026)
by: He, Langzhou, et al.
Published: (2026)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
by: Kajare, Prajwal Vijay, et al.
Published: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
by: Shravan, Rohan
Published: (2026)
by: Shravan, Rohan
Published: (2026)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
by: Keeman, Michael
Published: (2026)
by: Keeman, Michael
Published: (2026)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
by: Ubukata, Shunsuke
Published: (2026)
by: Ubukata, Shunsuke
Published: (2026)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
by: Walker, Nicholas
Published: (2024)
by: Walker, Nicholas
Published: (2024)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
On the Influence of Discourse Relations in Persuasive Texts
by: Turk, Nawar, et al.
Published: (2025)
by: Turk, Nawar, et al.
Published: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
by: Han, Xudong, et al.
Published: (2025)
by: Han, Xudong, et al.
Published: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
by: Bertina, Abbas, et al.
Published: (2025)
by: Bertina, Abbas, et al.
Published: (2025)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
by: Li, Zilin, et al.
Published: (2026)
by: Li, Zilin, et al.
Published: (2026)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
by: Pratyush, Spandan
Published: (2026)
by: Pratyush, Spandan
Published: (2026)
Continuous-Depth Transformers with Learned Control Dynamics
by: Jemley, Peter
Published: (2026)
by: Jemley, Peter
Published: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
by: He, Yanjin, et al.
Published: (2025)
by: He, Yanjin, et al.
Published: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
PersonalLLM: Tailoring LLMs to Individual Preferences
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
by: Vera, Patricio
Published: (2026)
by: Vera, Patricio
Published: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
by: Zhou, Tianyang, et al.
Published: (2026)
by: Zhou, Tianyang, et al.
Published: (2026)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
by: Young, Richard J., et al.
Published: (2025)
by: Young, Richard J., et al.
Published: (2025)
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
by: Sauter, Andreas, et al.
Published: (2026)
by: Sauter, Andreas, et al.
Published: (2026)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
Similar Items
-
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025) -
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026) -
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026) -
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
by: Zhang, Luyan
Published: (2025) -
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)