Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Qingru, Yu, Xiaodong, Singh, Chandan, Liu, Xiaodong, Liu, Liyuan, Gao, Jianfeng, Zhao, Tuo, Roth, Dan, Cheng, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
von: Zhang, Qingru, et al.
Veröffentlicht: (2023)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
Vector-ICL: In-context Learning with Continuous Vector Representations
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Learning a Decision Tree Algorithm with Transformers
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
Training Large Reasoning Models Efficiently via Progressive Thought Encoding
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
Test-time Recursive Thinking: Self-Improvement without External Feedback
von: Zhuang, Yufan, et al.
Veröffentlicht: (2026)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2026)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
SAS: Simulated Attention Score
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
DefSent+: Improving sentence embeddings of language models by projecting definition sentences into a quasi-isotropic or isotropic vector space of unlimited dictionary entries
von: Liu, Xiaodong
Veröffentlicht: (2024)
von: Liu, Xiaodong
Veröffentlicht: (2024)
Detoxification for LLM: From Dataset Itself
von: Shao, Wei, et al.
Veröffentlicht: (2026)
von: Shao, Wei, et al.
Veröffentlicht: (2026)
Language Models as Inductive Reasoners
von: Yang, Zonglin, et al.
Veröffentlicht: (2022)
von: Yang, Zonglin, et al.
Veröffentlicht: (2022)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
von: Chen, Tong, et al.
Veröffentlicht: (2024)
von: Chen, Tong, et al.
Veröffentlicht: (2024)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
GeoSteer: Faithful Chain-of-Thought Steering via Latent Manifold Gradients
von: Kazama, Kentaro, et al.
Veröffentlicht: (2026)
von: Kazama, Kentaro, et al.
Veröffentlicht: (2026)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
von: Feng, Mingqian, et al.
Veröffentlicht: (2026)
von: Feng, Mingqian, et al.
Veröffentlicht: (2026)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
von: Song, Jialin, et al.
Veröffentlicht: (2026)
von: Song, Jialin, et al.
Veröffentlicht: (2026)
ReasonAgain: Using Extractable Symbolic Programs to Evaluate Mathematical Reasoning
von: Yu, Xiaodong, et al.
Veröffentlicht: (2024)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2024)
Faithful Bi-Directional Model Steering via Distribution Matching and Distributed Interchange Interventions
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
Predicting Where Steering Vectors Succeed
von: Billa, Jayadev
Veröffentlicht: (2026)
von: Billa, Jayadev
Veröffentlicht: (2026)
Evaluating LLMs on Chinese Topic Constructions: A Research Proposal Inspired by Tian et al. (2024)
von: Yang, Xiaodong
Veröffentlicht: (2025)
von: Yang, Xiaodong
Veröffentlicht: (2025)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
Conflicts in Texts: Data, Implications and Challenges
von: Liu, Siyi, et al.
Veröffentlicht: (2025)
von: Liu, Siyi, et al.
Veröffentlicht: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
Knocking-Heads Attention
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanchao, et al.
Veröffentlicht: (2025)
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
von: Park, Keunhyeung, et al.
Veröffentlicht: (2025)
von: Park, Keunhyeung, et al.
Veröffentlicht: (2025)
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers
von: Gong, Linyuan, et al.
Veröffentlicht: (2023)
von: Gong, Linyuan, et al.
Veröffentlicht: (2023)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
von: Wang, Zheng, et al.
Veröffentlicht: (2024)
von: Wang, Zheng, et al.
Veröffentlicht: (2024)
Routesplain: Towards Faithful and Intervenable Routing for Software-related Tasks
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
von: Štorek, Adam, et al.
Veröffentlicht: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
von: Wang, Zezhong, et al.
Veröffentlicht: (2023)
Rethinking Interpretability in the Era of Large Language Models
von: Singh, Chandan, et al.
Veröffentlicht: (2024)
von: Singh, Chandan, et al.
Veröffentlicht: (2024)
Does a Global Perspective Help Prune Sparse MoEs Elegantly?
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2024)
von: Ben-Artzy, Amit, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
von: Zhang, Qingru, et al.
Veröffentlicht: (2023) -
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023) -
Vector-ICL: In-context Learning with Continuous Vector Representations
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024) -
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025) -
Learning a Decision Tree Algorithm with Transformers
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)