Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qingru, Singh, Chandan, Liu, Liyuan, Liu, Xiaodong, Yu, Bin, Gao, Jianfeng, Zhao, Tuo |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
by: Zhang, Qingru, et al.
Published: (2024)
by: Zhang, Qingru, et al.
Published: (2024)
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023)
by: Ge, Suyu, et al.
Published: (2023)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
Vector-ICL: In-context Learning with Continuous Vector Representations
by: Zhuang, Yufan, et al.
Published: (2024)
by: Zhuang, Yufan, et al.
Published: (2024)
Text Generation Beyond Discrete Token Sampling
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024)
by: Lopardo, Gianluigi, et al.
Published: (2024)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Rethinking Interpretability in the Era of Large Language Models
by: Singh, Chandan, et al.
Published: (2024)
by: Singh, Chandan, et al.
Published: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
Predicting Where Steering Vectors Succeed
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Harnessing Large Language Models as Post-hoc Correctors
by: Zhong, Zhiqiang, et al.
Published: (2024)
by: Zhong, Zhiqiang, et al.
Published: (2024)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
by: Hayou, Soufiane, et al.
Published: (2025)
by: Hayou, Soufiane, et al.
Published: (2025)
Crafting Interpretable Embeddings by Asking LLMs Questions
by: Benara, Vinamra, et al.
Published: (2024)
by: Benara, Vinamra, et al.
Published: (2024)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
SteerConf: Steering LLMs for Confidence Elicitation
by: Zhou, Ziang, et al.
Published: (2025)
by: Zhou, Ziang, et al.
Published: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
by: Heyen, Henning, et al.
Published: (2024)
by: Heyen, Henning, et al.
Published: (2024)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Test-time Recursive Thinking: Self-Improvement without External Feedback
by: Zhuang, Yufan, et al.
Published: (2026)
by: Zhuang, Yufan, et al.
Published: (2026)
Conformalized Exceptional Model Mining: Telling Where Your Model Performs (Not) Well
by: Du, Xin, et al.
Published: (2025)
by: Du, Xin, et al.
Published: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
by: Wang, Zheng, et al.
Published: (2024)
by: Wang, Zheng, et al.
Published: (2024)
TRA: Better Length Generalisation with Threshold Relative Attention
by: Opper, Mattia, et al.
Published: (2025)
by: Opper, Mattia, et al.
Published: (2025)
RoseRAG: Robust Retrieval-augmented Generation with Small-scale LLMs via Margin-aware Preference Optimization
by: Liu, Tianci, et al.
Published: (2025)
by: Liu, Tianci, et al.
Published: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)
by: Stickland, Asa Cooper, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Steer2Adapt: Dynamically Composing Steering Vectors Elicits Efficient Adaptation of LLMs
by: Han, Pengrui, et al.
Published: (2026)
by: Han, Pengrui, et al.
Published: (2026)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025)
by: Wang, Linlin, et al.
Published: (2025)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
by: Singh, Chandan, et al.
Published: (2026)
by: Singh, Chandan, et al.
Published: (2026)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
by: Zhang, Qingru, et al.
Published: (2025)
by: Zhang, Qingru, et al.
Published: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
Similar Items
-
Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
by: Zhang, Qingru, et al.
Published: (2024) -
Learning a Decision Tree Algorithm with Transformers
by: Zhuang, Yufan, et al.
Published: (2024) -
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024) -
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023) -
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)