Resonance RoPE: Improving Context Length Generalization of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Suyuchen, Kobyzev, Ivan, Lu, Peng, Rezagholizadeh, Mehdi, Liu, Bang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Periodic RoPE for Infinite Context LLMs
di: Huo, Simin
Pubblicazione: (2026)
di: Huo, Simin
Pubblicazione: (2026)
Scaling Laws of RoPE-based Extrapolation
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
di: Liu, Xiaoran, et al.
Pubblicazione: (2023)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
di: Pan, Dayan, et al.
Pubblicazione: (2025)
di: Pan, Dayan, et al.
Pubblicazione: (2025)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
di: Du, Yufeng, et al.
Pubblicazione: (2026)
di: Du, Yufeng, et al.
Pubblicazione: (2026)
Base of RoPE Bounds Context Length
di: Men, Xin, et al.
Pubblicazione: (2024)
di: Men, Xin, et al.
Pubblicazione: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
di: Li, Haoran, et al.
Pubblicazione: (2026)
di: Li, Haoran, et al.
Pubblicazione: (2026)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
di: Yang, Ning, et al.
Pubblicazione: (2026)
di: Yang, Ning, et al.
Pubblicazione: (2026)
Demystifying the Slash Pattern in Attention: The Role of RoPE
di: Cheng, Yuan, et al.
Pubblicazione: (2026)
di: Cheng, Yuan, et al.
Pubblicazione: (2026)
On the token distance modeling ability of higher RoPE attention dimension
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
di: Liu, Feilong
Pubblicazione: (2026)
di: Liu, Feilong
Pubblicazione: (2026)
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
di: Wang, Suyuchen, et al.
Pubblicazione: (2025)
di: Wang, Suyuchen, et al.
Pubblicazione: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
R$^3$Mem: Bridging Memory Retention and Retrieval via Reversible Compression
di: Wang, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Wang, Xiaoqiang, et al.
Pubblicazione: (2025)
ReGLA: Refining Gated Linear Attention
di: Lu, Peng, et al.
Pubblicazione: (2025)
di: Lu, Peng, et al.
Pubblicazione: (2025)
FACT: Examining the Effectiveness of Iterative Context Rewriting for Multi-fact Retrieval
di: Wang, Jinlin, et al.
Pubblicazione: (2024)
di: Wang, Jinlin, et al.
Pubblicazione: (2024)
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
di: Picov, Isaac, et al.
Pubblicazione: (2026)
di: Picov, Isaac, et al.
Pubblicazione: (2026)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
di: Lu, Peng, et al.
Pubblicazione: (2023)
di: Lu, Peng, et al.
Pubblicazione: (2023)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
di: Metel, Michael R., et al.
Pubblicazione: (2024)
di: Metel, Michael R., et al.
Pubblicazione: (2024)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
di: Li, Zichong, et al.
Pubblicazione: (2026)
di: Li, Zichong, et al.
Pubblicazione: (2026)
Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization
di: Urrutia, Felipe, et al.
Pubblicazione: (2026)
di: Urrutia, Felipe, et al.
Pubblicazione: (2026)
A Circular Argument : Does RoPE need to be Equivariant for Vision?
di: van de Geijn, Chase, et al.
Pubblicazione: (2025)
di: van de Geijn, Chase, et al.
Pubblicazione: (2025)
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
di: Han, Chi, et al.
Pubblicazione: (2023)
di: Han, Chi, et al.
Pubblicazione: (2023)
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
di: Zhao, Liang, et al.
Pubblicazione: (2024)
di: Zhao, Liang, et al.
Pubblicazione: (2024)
Frayed RoPE and Long Inputs: A Geometric Perspective
di: Wertheimer, Davis, et al.
Pubblicazione: (2026)
di: Wertheimer, Davis, et al.
Pubblicazione: (2026)
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
di: Yan, Yuchen, et al.
Pubblicazione: (2025)
di: Yan, Yuchen, et al.
Pubblicazione: (2025)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
di: Wang, Haonan, et al.
Pubblicazione: (2024)
di: Wang, Haonan, et al.
Pubblicazione: (2024)
When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models
di: Zheng, Yingming, et al.
Pubblicazione: (2025)
di: Zheng, Yingming, et al.
Pubblicazione: (2025)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
di: Wang, Xindi, et al.
Pubblicazione: (2024)
di: Wang, Xindi, et al.
Pubblicazione: (2024)
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
di: Gokmen, Ahmet Berke, et al.
Pubblicazione: (2025)
di: Gokmen, Ahmet Berke, et al.
Pubblicazione: (2025)
RoCar: A Relationship Network-based Evaluation Method for Large Language Models
di: Wang, Ming, et al.
Pubblicazione: (2023)
di: Wang, Ming, et al.
Pubblicazione: (2023)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
di: He, Guangxin, et al.
Pubblicazione: (2025)
di: He, Guangxin, et al.
Pubblicazione: (2025)
Genshin: General Shield for Natural Language Processing with Large Language Models
di: Peng, Xiao, et al.
Pubblicazione: (2024)
di: Peng, Xiao, et al.
Pubblicazione: (2024)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
di: An, Bang, et al.
Pubblicazione: (2025)
di: An, Bang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Periodic RoPE for Infinite Context LLMs
di: Huo, Simin
Pubblicazione: (2026) -
Scaling Laws of RoPE-based Extrapolation
di: Liu, Xiaoran, et al.
Pubblicazione: (2023) -
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
di: Pan, Dayan, et al.
Pubblicazione: (2025) -
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
di: Du, Yufeng, et al.
Pubblicazione: (2026) -
Base of RoPE Bounds Context Length
di: Men, Xin, et al.
Pubblicazione: (2024)