CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Seonglae, Wu, Zekun, Koshiyama, Adriano |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026)
by: Patel, Het, et al.
Published: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
by: Cui, Sasha, et al.
Published: (2025)
by: Cui, Sasha, et al.
Published: (2025)
Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
by: Liu, Ming
Published: (2026)
by: Liu, Ming
Published: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
by: Yang, Jisoo, et al.
Published: (2026)
by: Yang, Jisoo, et al.
Published: (2026)
Fuzzy, Symbolic, and Contextual: Enhancing LLM Instruction via Cognitive Scaffolding
by: Figueiredo, Vanessa
Published: (2025)
by: Figueiredo, Vanessa
Published: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
by: Liu, Zhongxin, et al.
Published: (2025)
by: Liu, Zhongxin, et al.
Published: (2025)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
by: Sun, Kangkang, et al.
Published: (2026)
by: Sun, Kangkang, et al.
Published: (2026)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
by: Agarwal, Mahak, et al.
Published: (2025)
by: Agarwal, Mahak, et al.
Published: (2025)
Eyla: Toward an Identity-Anchored LLM Architecture with Integrated Biological Priors -- Vision, Implementation Attempt, and Lessons from AI-Assisted Development
by: Aditto, Arif
Published: (2026)
by: Aditto, Arif
Published: (2026)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
by: Zhou, Yangyang, et al.
Published: (2026)
by: Zhou, Yangyang, et al.
Published: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
by: Resck, Lucas, et al.
Published: (2026)
by: Resck, Lucas, et al.
Published: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
by: Liao, Jianxing, et al.
Published: (2025)
by: Liao, Jianxing, et al.
Published: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
by: Cho, Hanjun, et al.
Published: (2026)
by: Cho, Hanjun, et al.
Published: (2026)
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
by: Patarlapalli, Sai Babu, et al.
Published: (2026)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
by: Pratyush, Spandan
Published: (2026)
by: Pratyush, Spandan
Published: (2026)
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification
by: Hilasaca, Kenji, et al.
Published: (2026)
by: Hilasaca, Kenji, et al.
Published: (2026)
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law
by: Shportun, Nazarii
Published: (2026)
by: Shportun, Nazarii
Published: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
by: Vera, Patricio
Published: (2026)
by: Vera, Patricio
Published: (2026)
Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
by: Keeman, Michael
Published: (2026)
by: Keeman, Michael
Published: (2026)
KAConvText: Novel Approach to Burmese Sentence Classification using Kolmogorov-Arnold Convolution
by: Thu, Ye Kyaw, et al.
Published: (2025)
by: Thu, Ye Kyaw, et al.
Published: (2025)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication
by: Zhao, Zhilong, et al.
Published: (2025)
by: Zhao, Zhilong, et al.
Published: (2025)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
by: Bouchekif, Abdessalam, et al.
Published: (2025)
by: Bouchekif, Abdessalam, et al.
Published: (2025)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Truth as a Compression Artifact in Language Model Training
by: Krestnikov, Konstantin
Published: (2026)
by: Krestnikov, Konstantin
Published: (2026)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
by: Zhang, Yuxuan
Published: (2025)
by: Zhang, Yuxuan
Published: (2025)
Similar Items
-
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026) -
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026) -
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
by: Patel, Het, et al.
Published: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025) -
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
by: Cui, Sasha, et al.
Published: (2025)