The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Chiang, Ting-Rui, Yogatama, Dani |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
Selective Rotary Position Embedding
by: Movahedi, Sajad, et al.
Published: (2025)
by: Movahedi, Sajad, et al.
Published: (2025)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
On Retrieval Augmentation and the Limitations of Language Model Training
by: Chiang, Ting-Rui, et al.
Published: (2023)
by: Chiang, Ting-Rui, et al.
Published: (2023)
Foundation Model's Embedded Representations May Detect Distribution Shift
by: Vargas, Max, et al.
Published: (2023)
by: Vargas, Max, et al.
Published: (2023)
LongEmbed: Extending Embedding Models for Long Context Retrieval
by: Zhu, Dawei, et al.
Published: (2024)
by: Zhu, Dawei, et al.
Published: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
Sparse Upcycling: Inference Inefficient Finetuning
by: Doubov, Sasha, et al.
Published: (2024)
by: Doubov, Sasha, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
by: Zhu, Shiyi, et al.
Published: (2023)
by: Zhu, Shiyi, et al.
Published: (2023)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Rotary Offset Features in Large Language Models
by: Jonasson, André
Published: (2025)
by: Jonasson, André
Published: (2025)
Measuring Intrinsic Dimension of Token Embeddings
by: Kataiwa, Takuya, et al.
Published: (2025)
by: Kataiwa, Takuya, et al.
Published: (2025)
Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
by: Awadhiya, Kanishk
Published: (2026)
by: Awadhiya, Kanishk
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
Improving Rare Word Translation With Dictionaries and Attention Masking
by: Sible, Kenneth J., et al.
Published: (2024)
by: Sible, Kenneth J., et al.
Published: (2024)
Efficient Matrix Implementation for Rotary Position Embedding
by: Minqi, Chen, et al.
Published: (2026)
by: Minqi, Chen, et al.
Published: (2026)
Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns
by: Zhao, Yang, et al.
Published: (2024)
by: Zhao, Yang, et al.
Published: (2024)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
by: Lee, Changhun, et al.
Published: (2025)
by: Lee, Changhun, et al.
Published: (2025)
Retrieval Backward Attention without Additional Training: Enhance Embeddings of Large Language Models via Repetition
by: Duan, Yifei, et al.
Published: (2025)
by: Duan, Yifei, et al.
Published: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Simulating Hard Attention Using Soft Attention
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
by: Sarkar, Soumajyoti, et al.
Published: (2024)
by: Sarkar, Soumajyoti, et al.
Published: (2024)
Context-aware Rotary Position Embedding
by: Veisi, Ali, et al.
Published: (2025)
by: Veisi, Ali, et al.
Published: (2025)
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
by: Yoon, Jinsung, et al.
Published: (2024)
by: Yoon, Jinsung, et al.
Published: (2024)
MrRoPE: Mixed-radix Rotary Position Embedding
by: Tian, Qingyuan, et al.
Published: (2026)
by: Tian, Qingyuan, et al.
Published: (2026)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
ProxyAttn: Guided Sparse Attention via Representative Heads
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
MAGE: Multi-Head Attention Guided Embeddings for Low Resource Sentiment Classification
by: Vashisht, Varun, et al.
Published: (2025)
by: Vashisht, Varun, et al.
Published: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
by: Hsieh, Cheng-Yu, et al.
Published: (2024)
Learning Word Embedding with Better Distance Weighting and Window Size Scheduling
by: Yang, Chaohao, et al.
Published: (2024)
by: Yang, Chaohao, et al.
Published: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
Similar Items
-
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
by: Chiang, Ting-Rui, et al.
Published: (2024) -
Selective Rotary Position Embedding
by: Movahedi, Sajad, et al.
Published: (2025) -
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024) -
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026) -
On Retrieval Augmentation and the Limitations of Language Model Training
by: Chiang, Ting-Rui, et al.
Published: (2023)