A Circular Argument : Does RoPE need to be Equivariant for Vision?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | van de Geijn, Chase, Lüddecke, Timo, Turishcheva, Polina, Ecker, Alexander S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025)
Do traveling waves make good positional encodings?
von: van de Geijn, Chase, et al.
Veröffentlicht: (2025)
von: van de Geijn, Chase, et al.
Veröffentlicht: (2025)
ReRoPE: Repurposing RoPE for Relative Camera Control
von: Li, Chunyang, et al.
Veröffentlicht: (2026)
von: Li, Chunyang, et al.
Veröffentlicht: (2026)
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
von: Roberts, Jonathan, et al.
Veröffentlicht: (2023)
von: Roberts, Jonathan, et al.
Veröffentlicht: (2023)
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
von: Schenck, Connor, et al.
Veröffentlicht: (2025)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
von: Liu, Haoyu, et al.
Veröffentlicht: (2026)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
von: Li, Bozhou, et al.
Veröffentlicht: (2025)
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
von: Kung, Gustavo Chau Loo, et al.
Veröffentlicht: (2026)
von: Kung, Gustavo Chau Loo, et al.
Veröffentlicht: (2026)
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
von: Yan, Feihong, et al.
Veröffentlicht: (2026)
von: Yan, Feihong, et al.
Veröffentlicht: (2026)
Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
von: Yesiltepe, Hidir, et al.
Veröffentlicht: (2025)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
von: Liu, Yuxi, et al.
Veröffentlicht: (2026)
von: Liu, Yuxi, et al.
Veröffentlicht: (2026)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
von: Wei, Tianyi, et al.
Veröffentlicht: (2025)
von: Wei, Tianyi, et al.
Veröffentlicht: (2025)
Relaxed Rotational Equivariance via $G$-Biases in Vision
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2024)
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
Periodic RoPE for Infinite Context LLMs
von: Huo, Simin
Veröffentlicht: (2026)
von: Huo, Simin
Veröffentlicht: (2026)
Scaling Laws of RoPE-based Extrapolation
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2023)
Zero-Shot Multi-Animal Tracking in the Wild
von: Meier, Jan Frederik, et al.
Veröffentlicht: (2025)
von: Meier, Jan Frederik, et al.
Veröffentlicht: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
von: Nordström, David, et al.
Veröffentlicht: (2025)
von: Nordström, David, et al.
Veröffentlicht: (2025)
Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
von: Cedro, Mateusz, et al.
Veröffentlicht: (2026)
RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Fengxiang, et al.
Veröffentlicht: (2025)
RoGA: Towards Generalizable Deepfake Detection through Robust Gradient Alignment
von: Qiu, Lingyu, et al.
Veröffentlicht: (2025)
von: Qiu, Lingyu, et al.
Veröffentlicht: (2025)
Equivariant Flow Matching for Point Cloud Assembly
von: Wang, Ziming, et al.
Veröffentlicht: (2025)
von: Wang, Ziming, et al.
Veröffentlicht: (2025)
RoSe: Robust Self-supervised Stereo Matching under Adverse Weather Conditions
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
RoBus: A Multimodal Dataset for Controllable Road Networks and Building Layouts Generation
von: Li, Tao, et al.
Veröffentlicht: (2024)
von: Li, Tao, et al.
Veröffentlicht: (2024)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
von: Park, Seulki, et al.
Veröffentlicht: (2023)
von: Park, Seulki, et al.
Veröffentlicht: (2023)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
von: Jiao, Pengkun, et al.
Veröffentlicht: (2024)
Vision Transformers: the threat of realistic adversarial patches
von: Cools, Kasper, et al.
Veröffentlicht: (2025)
von: Cools, Kasper, et al.
Veröffentlicht: (2025)
Few-shot Implicit Function Generation via Equivariance
von: Huang, Suizhi, et al.
Veröffentlicht: (2025)
von: Huang, Suizhi, et al.
Veröffentlicht: (2025)
Learning Generalizable Shape Completion with SIM(3) Equivariance
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
von: Saied, Youssef, et al.
Veröffentlicht: (2026)
von: Saied, Youssef, et al.
Veröffentlicht: (2026)
Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis
von: Tong, Ran, et al.
Veröffentlicht: (2025)
von: Tong, Ran, et al.
Veröffentlicht: (2025)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
von: He, Jingtao, et al.
Veröffentlicht: (2026)
von: He, Jingtao, et al.
Veröffentlicht: (2026)
GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
von: Yao, Yupu, et al.
Veröffentlicht: (2025)
von: Yao, Yupu, et al.
Veröffentlicht: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
von: Cheng, Yuan, et al.
Veröffentlicht: (2026)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
A Simple Efficiency Incremental Learning Framework via Vision-Language Model with Nonlinear Multi-Adapters
von: Luo, Haihua, et al.
Veröffentlicht: (2026)
von: Luo, Haihua, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers
von: Gokmen, Ahmet Berke, et al.
Veröffentlicht: (2025) -
Do traveling waves make good positional encodings?
von: van de Geijn, Chase, et al.
Veröffentlicht: (2025) -
ReRoPE: Repurposing RoPE for Relative Camera Control
von: Li, Chunyang, et al.
Veröffentlicht: (2026) -
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
von: Roberts, Jonathan, et al.
Veröffentlicht: (2023) -
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
von: Schenck, Connor, et al.
Veröffentlicht: (2025)