Which Attention Heads Matter for In-Context Learning?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Kayo, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Long Chain-of-Thought Reasoning Across Languages
von: Barua, Josh, et al.
Veröffentlicht: (2025)
von: Barua, Josh, et al.
Veröffentlicht: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026)
von: Svete, Anej, et al.
Veröffentlicht: (2026)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024)
von: Collins, Liam, et al.
Veröffentlicht: (2024)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Decomposing Attention To Find Context-Sensitive Neurons
von: Gibson, Alex
Veröffentlicht: (2025)
von: Gibson, Alex
Veröffentlicht: (2025)
Toward In-Context Teaching: Adapting Examples to Students' Misconceptions
von: Ross, Alexis, et al.
Veröffentlicht: (2024)
von: Ross, Alexis, et al.
Veröffentlicht: (2024)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
von: Lee, Changhun, et al.
Veröffentlicht: (2025)
von: Lee, Changhun, et al.
Veröffentlicht: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
von: Chen, Yingfa, et al.
Veröffentlicht: (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
HiCI: Hierarchical Construction-Integration for Long-Context Attention
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
Is In-Context Learning Learning?
von: de Wynter, Adrian
Veröffentlicht: (2025)
von: de Wynter, Adrian
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
How Private is Your Attention? Bridging Privacy with In-Context Learning
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
von: Bonnerjee, Soham, et al.
Veröffentlicht: (2025)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2024)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2024)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
von: Sehanobish, Arijit, et al.
Veröffentlicht: (2026)
von: Sehanobish, Arijit, et al.
Veröffentlicht: (2026)
Memorization in In-Context Learning
von: Golchin, Shahriar, et al.
Veröffentlicht: (2024)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2024)
Revisiting In-Context Learning with Long Context Language Models
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
von: Baek, Jinheon, et al.
Veröffentlicht: (2024)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025) -
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025) -
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)