Selective Rotary Position Embedding
Fuente:
arXiv
Saved in:
| Main Authors: | Movahedi, Sajad, Carstensen, Timur, Afzal, Arshia, Hutter, Frank, Orvieto, Antonio, Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
by: Siems, Julien, et al.
Published: (2025)
by: Siems, Julien, et al.
Published: (2025)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)
by: Afzal, Arshia, et al.
Published: (2024)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
by: Movahedi, Sajad, et al.
Published: (2024)
by: Movahedi, Sajad, et al.
Published: (2024)
Fixed-Point RNNs: Interpolating from Diagonal to Dense
by: Movahedi, Sajad, et al.
Published: (2025)
by: Movahedi, Sajad, et al.
Published: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
by: Chiang, Ting-Rui, et al.
Published: (2025)
by: Chiang, Ting-Rui, et al.
Published: (2025)
Quickly Tuning Foundation Models for Image Segmentation
by: Das, Breenda, et al.
Published: (2025)
by: Das, Breenda, et al.
Published: (2025)
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
by: Sarkar, Soumajyoti, et al.
Published: (2024)
by: Sarkar, Soumajyoti, et al.
Published: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
by: Carstensen, Timur, et al.
Published: (2025)
by: Carstensen, Timur, et al.
Published: (2025)
Unlearning as multi-task optimization: A normalized gradient difference approach with an adaptive learning rate
by: Bu, Zhiqi, et al.
Published: (2024)
by: Bu, Zhiqi, et al.
Published: (2024)
Revisiting Character-level Adversarial Attacks for Language Models
by: Rocamora, Elias Abad, et al.
Published: (2024)
by: Rocamora, Elias Abad, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
by: Candogan, Leyla Naz, et al.
Published: (2025)
by: Candogan, Leyla Naz, et al.
Published: (2025)
Rotary Offset Features in Large Language Models
by: Jonasson, André
Published: (2025)
by: Jonasson, André
Published: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Efficient Matrix Implementation for Rotary Position Embedding
by: Minqi, Chen, et al.
Published: (2026)
by: Minqi, Chen, et al.
Published: (2026)
SemEval-2025 Task 4: Unlearning sensitive content from Large Language Models
by: Ramakrishna, Anil, et al.
Published: (2025)
by: Ramakrishna, Anil, et al.
Published: (2025)
LUME: LLM Unlearning with Multitask Evaluations
by: Ramakrishna, Anil, et al.
Published: (2025)
by: Ramakrishna, Anil, et al.
Published: (2025)
TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
by: Moroshan, Vladyslav, et al.
Published: (2025)
by: Moroshan, Vladyslav, et al.
Published: (2025)
Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
by: Nezhurina, Marianna, et al.
Published: (2025)
by: Nezhurina, Marianna, et al.
Published: (2025)
Context-aware Rotary Position Embedding
by: Veisi, Ali, et al.
Published: (2025)
by: Veisi, Ali, et al.
Published: (2025)
Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
by: Spiegelhalter, Urs, et al.
Published: (2025)
by: Spiegelhalter, Urs, et al.
Published: (2025)
Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
DoPE: Denoising Rotary Position Embedding
by: Xiong, Jing, et al.
Published: (2025)
by: Xiong, Jing, et al.
Published: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
Membership Inference Attacks against Large Vision-Language Models
by: Li, Zhan, et al.
Published: (2024)
by: Li, Zhan, et al.
Published: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Generalized Interpolating Discrete Diffusion
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings
by: Rambaud, Victor, et al.
Published: (2025)
by: Rambaud, Victor, et al.
Published: (2025)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
by: Bergerault, Antoine, et al.
Published: (2026)
by: Bergerault, Antoine, et al.
Published: (2026)
Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
by: Krishnakumar, Arjun, et al.
Published: (2025)
by: Krishnakumar, Arjun, et al.
Published: (2025)
CRoPE: Efficient Parametrization of Rotary Positional Embedding
by: Lou, Beicheng, et al.
Published: (2026)
by: Lou, Beicheng, et al.
Published: (2026)
Similar Items
-
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026) -
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025) -
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025) -
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
by: Siems, Julien, et al.
Published: (2025) -
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)