Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Haolin, Cho, Hakaze, Zhong, Yiqiao, Inoue, Naoya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
StaICC: Standardized Evaluation for Classification Task in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Mechanism of Task-oriented Information Removal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
by: Kato, Mariko, et al.
Published: (2025)
by: Kato, Mariko, et al.
Published: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Revisiting In-context Learning Inference Circuit in Large Language Models
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Mechanistic Fine-tuning for In-context Learning
by: Cho, Hakaze, et al.
Published: (2025)
by: Cho, Hakaze, et al.
Published: (2025)
Understanding Token Probability Encoding in Output Embeddings
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Token-based Decision Criteria Are Suboptimal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
Measuring Intrinsic Dimension of Token Embeddings
by: Kataiwa, Takuya, et al.
Published: (2025)
by: Kataiwa, Takuya, et al.
Published: (2025)
How does Multi-Task Training Affect Transformer In-Context Capabilities? Investigations with Function Classes
by: Bhasin, Harmon, et al.
Published: (2024)
by: Bhasin, Harmon, et al.
Published: (2024)
Label Words as Local Task Vectors in In-Context Learning
by: Zheng, Bowen, et al.
Published: (2024)
by: Zheng, Bowen, et al.
Published: (2024)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
by: He, Xingyang, et al.
Published: (2025)
by: He, Xingyang, et al.
Published: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
by: Liu, Siran, et al.
Published: (2026)
by: Liu, Siran, et al.
Published: (2026)
One Task Vector is not Enough: A Large-Scale Study for In-Context Learning
by: Tikhonov, Pavel, et al.
Published: (2025)
by: Tikhonov, Pavel, et al.
Published: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
In-Context Learning State Vector with Inner and Momentum Optimization
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
NoisyICL: A Little Noise in Model Parameters Calibrates In-context Learning
by: Zhao, Yufeng, et al.
Published: (2024)
by: Zhao, Yufeng, et al.
Published: (2024)
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
by: Kwon, Jihoon, et al.
Published: (2026)
by: Kwon, Jihoon, et al.
Published: (2026)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
DeCoVec: Building Decoding Space based Task Vector for Large Language Models via In-Context Learning
by: Li, Feiyang, et al.
Published: (2026)
by: Li, Feiyang, et al.
Published: (2026)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
Phase Diagram of Vision Large Language Models Inference: A Perspective from Interaction across Image and Instruction
by: Wei, Houjing, et al.
Published: (2024)
by: Wei, Houjing, et al.
Published: (2024)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
by: Song, Jiajun, et al.
Published: (2024)
by: Song, Jiajun, et al.
Published: (2024)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
by: Yang, Lijie, et al.
Published: (2025)
by: Yang, Lijie, et al.
Published: (2025)
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
by: Tezuka, Hinata, et al.
Published: (2025)
by: Tezuka, Hinata, et al.
Published: (2025)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
by: Ghaddar, Abbas, et al.
Published: (2026)
by: Ghaddar, Abbas, et al.
Published: (2026)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Knocking-Heads Attention
by: Zhou, Zhanchao, et al.
Published: (2025)
by: Zhou, Zhanchao, et al.
Published: (2025)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
Similar Items
-
Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis
by: Yang, Haolin, et al.
Published: (2025) -
Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight
by: Yang, Haolin, et al.
Published: (2025) -
StaICC: Standardized Evaluation for Classification Task in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025) -
Mechanism of Task-oriented Information Removal in In-context Learning
by: Cho, Hakaze, et al.
Published: (2025) -
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
by: Yan, Hao, et al.
Published: (2026)