A Circular Argument : Does RoPE need to be Equivariant for Vision?
Fuente:
arXiv
Saved in:
| Main Authors: | van de Geijn, Chase, Lüddecke, Timo, Turishcheva, Polina, Ecker, Alexander S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
by: Gokmen, Ahmet Berke, et al.
Published: (2025)
Do traveling waves make good positional encodings?
by: van de Geijn, Chase, et al.
Published: (2025)
by: van de Geijn, Chase, et al.
Published: (2025)
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026)
by: Li, Chunyang, et al.
Published: (2026)
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
by: Roberts, Jonathan, et al.
Published: (2023)
by: Roberts, Jonathan, et al.
Published: (2023)
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
by: Schenck, Connor, et al.
Published: (2025)
by: Schenck, Connor, et al.
Published: (2025)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
by: Mikaeili, Aryan, et al.
Published: (2026)
by: Mikaeili, Aryan, et al.
Published: (2026)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
by: Ahuja, Rahul, et al.
Published: (2026)
by: Ahuja, Rahul, et al.
Published: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
by: Yan, Feihong, et al.
Published: (2026)
by: Yan, Feihong, et al.
Published: (2026)
Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
by: Yesiltepe, Hidir, et al.
Published: (2025)
by: Yesiltepe, Hidir, et al.
Published: (2025)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
by: Wei, Tianyi, et al.
Published: (2025)
by: Wei, Tianyi, et al.
Published: (2025)
Relaxed Rotational Equivariance via $G$-Biases in Vision
by: Wu, Zhiqiang, et al.
Published: (2024)
by: Wu, Zhiqiang, et al.
Published: (2024)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
by: Yang, Yang, et al.
Published: (2026)
by: Yang, Yang, et al.
Published: (2026)
Periodic RoPE for Infinite Context LLMs
by: Huo, Simin
Published: (2026)
by: Huo, Simin
Published: (2026)
Scaling Laws of RoPE-based Extrapolation
by: Liu, Xiaoran, et al.
Published: (2023)
by: Liu, Xiaoran, et al.
Published: (2023)
Zero-Shot Multi-Animal Tracking in the Wild
by: Meier, Jan Frederik, et al.
Published: (2025)
by: Meier, Jan Frederik, et al.
Published: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
by: Cedro, Mateusz, et al.
Published: (2026)
by: Cedro, Mateusz, et al.
Published: (2026)
RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
RoGA: Towards Generalizable Deepfake Detection through Robust Gradient Alignment
by: Qiu, Lingyu, et al.
Published: (2025)
by: Qiu, Lingyu, et al.
Published: (2025)
Equivariant Flow Matching for Point Cloud Assembly
by: Wang, Ziming, et al.
Published: (2025)
by: Wang, Ziming, et al.
Published: (2025)
RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
RoBus: A Multimodal Dataset for Controllable Road Networks and Building Layouts Generation
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
by: Lin, Haokun, et al.
Published: (2024)
by: Lin, Haokun, et al.
Published: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
by: Park, Seulki, et al.
Published: (2023)
by: Park, Seulki, et al.
Published: (2023)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Vision Transformers: the threat of realistic adversarial patches
by: Cools, Kasper, et al.
Published: (2025)
by: Cools, Kasper, et al.
Published: (2025)
Few-shot Implicit Function Generation via Equivariance
by: Huang, Suizhi, et al.
Published: (2025)
by: Huang, Suizhi, et al.
Published: (2025)
Learning Generalizable Shape Completion with SIM(3) Equivariance
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
by: Saied, Youssef, et al.
Published: (2026)
by: Saied, Youssef, et al.
Published: (2026)
Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis
by: Tong, Ran, et al.
Published: (2025)
by: Tong, Ran, et al.
Published: (2025)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
by: He, Jingtao, et al.
Published: (2026)
by: He, Jingtao, et al.
Published: (2026)
GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
by: Yao, Yupu, et al.
Published: (2025)
by: Yao, Yupu, et al.
Published: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
by: Limpijankit, Marvin, et al.
Published: (2026)
by: Limpijankit, Marvin, et al.
Published: (2026)
A Simple Efficiency Incremental Learning Framework via Vision-Language Model with Nonlinear Multi-Adapters
by: Luo, Haihua, et al.
Published: (2026)
by: Luo, Haihua, et al.
Published: (2026)
Similar Items
-
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
by: Gokmen, Ahmet Berke, et al.
Published: (2025) -
Do traveling waves make good positional encodings?
by: van de Geijn, Chase, et al.
Published: (2025) -
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026) -
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
by: Roberts, Jonathan, et al.
Published: (2023) -
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
by: Schenck, Connor, et al.
Published: (2025)