Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Liu, Feilong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
von: Du, Yufeng, et al.
Veröffentlicht: (2026)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
von: Li, Haoran, et al.
Veröffentlicht: (2026)
von: Li, Haoran, et al.
Veröffentlicht: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Periodic RoPE for Infinite Context LLMs
von: Huo, Simin
Veröffentlicht: (2026)
von: Huo, Simin
Veröffentlicht: (2026)
Rethinking RoPE: A Mathematical Blueprint for N-dimensional Positional Embedding
von: Liu, Haiping, et al.
Veröffentlicht: (2025)
von: Liu, Haiping, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Frayed RoPE and Long Inputs: A Geometric Perspective
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
von: Wertheimer, Davis, et al.
Veröffentlicht: (2026)
Scaling Laws of RoPE-based Extrapolation
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
Base of RoPE Bounds Context Length
von: Men, Xin, et al.
Veröffentlicht: (2024)
von: Men, Xin, et al.
Veröffentlicht: (2024)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
von: Urrutia, Felipe, et al.
Veröffentlicht: (2026)
von: Urrutia, Felipe, et al.
Veröffentlicht: (2026)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
von: Zhu, Shiyi, et al.
Veröffentlicht: (2023)
von: Zhu, Shiyi, et al.
Veröffentlicht: (2023)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
von: Yang, Ning, et al.
Veröffentlicht: (2026)
von: Yang, Ning, et al.
Veröffentlicht: (2026)
On the token distance modeling ability of higher RoPE attention dimension
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks
von: Zhang, Yaobo
Veröffentlicht: (2026)
von: Zhang, Yaobo
Veröffentlicht: (2026)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2026)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
von: Liu, Feilong
Veröffentlicht: (2026)
von: Liu, Feilong
Veröffentlicht: (2026)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
von: Pan, Dayan, et al.
Veröffentlicht: (2025)
MrRoPE: Mixed-radix Rotary Position Embedding
von: Tian, Qingyuan, et al.
Veröffentlicht: (2026)
von: Tian, Qingyuan, et al.
Veröffentlicht: (2026)
Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models
von: Zhang, Chenyu, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2026)
Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
PoPE: Legendre Orthogonal Polynomials Based Position Encoding for Large Language Models
von: Aggarwal, Arpit
Veröffentlicht: (2024)
von: Aggarwal, Arpit
Veröffentlicht: (2024)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
Selective Rotary Position Embedding
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Meizhi, et al.
Veröffentlicht: (2024)
SG-UniBuc-NLP at SemEval-2026 Task 6: Multi-Head RoBERTa with Chunking for Long-Context Evasion Detection
von: Stefan, Gabriel, et al.
Veröffentlicht: (2026)
von: Stefan, Gabriel, et al.
Veröffentlicht: (2026)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis
von: Fu, Yao
Veröffentlicht: (2024)
von: Fu, Yao
Veröffentlicht: (2024)
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2024)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2024)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
von: Li, Zichong, et al.
Veröffentlicht: (2026)
von: Li, Zichong, et al.
Veröffentlicht: (2026)
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
von: Wang, Haonan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
von: Du, Yufeng, et al.
Veröffentlicht: (2026) -
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024) -
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
von: Li, Haoran, et al.
Veröffentlicht: (2026) -
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024) -
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)