Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Valentino, Marco, Kim, Geonhee, Dalal, Dhairya, Zhao, Zhixue, Freitas, André |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
Inference to the Best Explanation in Large Language Models
by: Dalal, Dhairya, et al.
Published: (2024)
by: Dalal, Dhairya, et al.
Published: (2024)
PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Reasoning with Natural Language Explanations
by: Valentino, Marco, et al.
Published: (2024)
by: Valentino, Marco, et al.
Published: (2024)
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations
by: Ranaldi, Leonardo, et al.
Published: (2024)
by: Ranaldi, Leonardo, et al.
Published: (2024)
Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models
by: Maraia, Gabriele, et al.
Published: (2026)
by: Maraia, Gabriele, et al.
Published: (2026)
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
by: Jullien, Mael, et al.
Published: (2025)
by: Jullien, Mael, et al.
Published: (2025)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials
by: Jullien, Mael, et al.
Published: (2024)
by: Jullien, Mael, et al.
Published: (2024)
A Differentiable Integer Linear Programming Solver for Explanation-Based Natural Language Inference
by: Thayaparan, Mokanarangan, et al.
Published: (2024)
by: Thayaparan, Mokanarangan, et al.
Published: (2024)
FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models
by: Weng, Zixuan, et al.
Published: (2026)
by: Weng, Zixuan, et al.
Published: (2026)
A Semantic Search Pipeline for Causality-driven Adhoc Information Retrieval
by: Dalal, Dhairya, et al.
Published: (2025)
by: Dalal, Dhairya, et al.
Published: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
by: Vasilakes, Jake, et al.
Published: (2025)
by: Vasilakes, Jake, et al.
Published: (2025)
Activation Scaling for Steering and Interpreting Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
by: Kabra, Sanchit, et al.
Published: (2025)
by: Kabra, Sanchit, et al.
Published: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
by: Chrysostomou, George, et al.
Published: (2023)
by: Chrysostomou, George, et al.
Published: (2023)
Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models
by: Tan, Xingwei, et al.
Published: (2026)
by: Tan, Xingwei, et al.
Published: (2026)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Integrating Expert Knowledge into Logical Programs via LLMs
by: Górski, Franciszek, et al.
Published: (2025)
by: Górski, Franciszek, et al.
Published: (2025)
Endogenous Resistance to Activation Steering in Language Models
by: McKenzie, Alex, et al.
Published: (2026)
by: McKenzie, Alex, et al.
Published: (2026)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)
by: Cheng, Yinjie, et al.
Published: (2025)
ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models
by: Guang, Jiahui, et al.
Published: (2026)
by: Guang, Jiahui, et al.
Published: (2026)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
by: Feng, Jiarui, et al.
Published: (2025)
by: Feng, Jiarui, et al.
Published: (2025)
SAKE: Steering Activations for Knowledge Editing
by: Scialanga, Marco, et al.
Published: (2025)
by: Scialanga, Marco, et al.
Published: (2025)
Model Fusion through Bayesian Optimization in Language Model Fine-Tuning
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya
by: Sathish, Sharath
Published: (2026)
by: Sathish, Sharath
Published: (2026)
On Effects of Steering Latent Representation for Large Language Model Unlearning
by: Huu-Tien, Dang, et al.
Published: (2024)
by: Huu-Tien, Dang, et al.
Published: (2024)
Mitigating Overthinking through Reasoning Shaping
by: Song, Feifan, et al.
Published: (2025)
by: Song, Feifan, et al.
Published: (2025)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps
by: Hsiao, Yen-Che, et al.
Published: (2025)
by: Hsiao, Yen-Che, et al.
Published: (2025)
Self-Steering Language Models
by: Grand, Gabriel, et al.
Published: (2025)
by: Grand, Gabriel, et al.
Published: (2025)
Multi-property Steering of Large Language Models with Dynamic Activation Composition
by: Scalena, Daniel, et al.
Published: (2024)
by: Scalena, Daniel, et al.
Published: (2024)
Similar Items
-
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024) -
Inference to the Best Explanation in Large Language Models
by: Dalal, Dhairya, et al.
Published: (2024) -
PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
by: Quan, Xin, et al.
Published: (2025) -
Reasoning with Natural Language Explanations
by: Valentino, Marco, et al.
Published: (2024) -
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations
by: Ranaldi, Leonardo, et al.
Published: (2024)