Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zichong, Liang, Chen, Ren, Liliang, Zhao, Tuo, Shen, Yelong, Chen, Weizhu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Base of RoPE Bounds Context Length
by: Men, Xin, et al.
Published: (2024)
by: Men, Xin, et al.
Published: (2024)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
by: Zhong, Meizhi, et al.
Published: (2024)
by: Zhong, Meizhi, et al.
Published: (2024)
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
by: Ren, Liliang, et al.
Published: (2024)
by: Ren, Liliang, et al.
Published: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
by: Li, Haoran, et al.
Published: (2026)
by: Li, Haoran, et al.
Published: (2026)
Periodic RoPE for Infinite Context LLMs
by: Huo, Simin
Published: (2026)
by: Huo, Simin
Published: (2026)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
by: Du, Yufeng, et al.
Published: (2026)
by: Du, Yufeng, et al.
Published: (2026)
LongRoPE2: Near-Lossless LLM Context Window Scaling
by: Shang, Ning, et al.
Published: (2025)
by: Shang, Ning, et al.
Published: (2025)
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
by: Wang, Haonan, et al.
Published: (2024)
by: Wang, Haonan, et al.
Published: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attribution
by: Picov, Isaac, et al.
Published: (2026)
by: Picov, Isaac, et al.
Published: (2026)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
Frayed RoPE and Long Inputs: A Geometric Perspective
by: Wertheimer, Davis, et al.
Published: (2026)
by: Wertheimer, Davis, et al.
Published: (2026)
NorMuon: Making Muon more efficient and scalable
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Scaling Laws of RoPE-based Extrapolation
by: Liu, Xiaoran, et al.
Published: (2023)
by: Liu, Xiaoran, et al.
Published: (2023)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
RoSA: Enhancing Parameter-Efficient Fine-Tuning via RoPE-aware Selective Adaptation in Large Language Models
by: Pan, Dayan, et al.
Published: (2025)
by: Pan, Dayan, et al.
Published: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
On the token distance modeling ability of higher RoPE attention dimension
by: Hong, Xiangyu, et al.
Published: (2024)
by: Hong, Xiangyu, et al.
Published: (2024)
LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
by: Ding, Yiran, et al.
Published: (2024)
by: Ding, Yiran, et al.
Published: (2024)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
by: Khan, Mohammad Aflah, et al.
Published: (2026)
by: Khan, Mohammad Aflah, et al.
Published: (2026)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
by: Ren, Liliang, et al.
Published: (2026)
by: Ren, Liliang, et al.
Published: (2026)
Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks
by: Zhang, Yaobo
Published: (2026)
by: Zhang, Yaobo
Published: (2026)
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026)
by: Li, Chunyang, et al.
Published: (2026)
LLMs Can Generate a Better Answer by Aggregating Their Own Responses
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
by: Ren, Liliang, et al.
Published: (2025)
by: Ren, Liliang, et al.
Published: (2025)
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
Test-time Recursive Thinking: Self-Improvement without External Feedback
by: Zhuang, Yufan, et al.
Published: (2026)
by: Zhuang, Yufan, et al.
Published: (2026)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
by: Zhang, Xinsen, et al.
Published: (2026)
by: Zhang, Xinsen, et al.
Published: (2026)
Exploring the Mystery of Influential Data for Mathematical Reasoning
by: Ni, Xinzhe, et al.
Published: (2024)
by: Ni, Xinzhe, et al.
Published: (2024)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
by: Qiao, Ye, et al.
Published: (2025)
by: Qiao, Ye, et al.
Published: (2025)
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
by: Yang, Ziyi, et al.
Published: (2025)
by: Yang, Ziyi, et al.
Published: (2025)
Similar Items
-
Base of RoPE Bounds Context Length
by: Men, Xin, et al.
Published: (2024) -
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
by: Zhong, Meizhi, et al.
Published: (2024) -
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
by: Ren, Liliang, et al.
Published: (2024) -
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
by: Li, Haoran, et al.
Published: (2026) -
Periodic RoPE for Infinite Context LLMs
by: Huo, Simin
Published: (2026)