RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yu, Jeon, Minsik, Chang, Jen-Hao Rick, Tuzel, Oncel, Tulsiani, Shubham |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026)
by: Cong, Zhongxiao, et al.
Published: (2026)
LiTo: Surface Light Field Tokenization
by: Chang, Jen-Hao Rick, et al.
Published: (2026)
by: Chang, Jen-Hao Rick, et al.
Published: (2026)
Velox: Learning Representations of 4D Geometry and Appearance
by: Malik, Anagh, et al.
Published: (2026)
by: Malik, Anagh, et al.
Published: (2026)
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024)
by: Zhao, Qitao, et al.
Published: (2024)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
Cameras as Rays: Pose Estimation via Ray Diffusion
by: Zhang, Jason Y., et al.
Published: (2024)
by: Zhang, Jason Y., et al.
Published: (2024)
LightSwitch: Multi-view Relighting with Material-guided Diffusion
by: Litman, Yehonathan, et al.
Published: (2025)
by: Litman, Yehonathan, et al.
Published: (2025)
3D Shape Tokenization via Latent Flow Matching
by: Chang, Jen-Hao Rick, et al.
Published: (2024)
by: Chang, Jen-Hao Rick, et al.
Published: (2024)
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
by: Jin, Seonghyun, et al.
Published: (2026)
by: Jin, Seonghyun, et al.
Published: (2026)
Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
by: Ahn, Byeongjoo, et al.
Published: (2023)
by: Ahn, Byeongjoo, et al.
Published: (2023)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
by: Ahuja, Rahul, et al.
Published: (2026)
by: Ahuja, Rahul, et al.
Published: (2026)
ReRoPE: Repurposing RoPE for Relative Camera Control
by: Li, Chunyang, et al.
Published: (2026)
by: Li, Chunyang, et al.
Published: (2026)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
by: Pay, Quan Bi, et al.
Published: (2025)
by: Pay, Quan Bi, et al.
Published: (2025)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
by: Mikaeili, Aryan, et al.
Published: (2026)
by: Mikaeili, Aryan, et al.
Published: (2026)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
Diverse Score Distillation
by: Xu, Yanbo, et al.
Published: (2024)
by: Xu, Yanbo, et al.
Published: (2024)
SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
by: Bokhovkin, Alexey, et al.
Published: (2024)
by: Bokhovkin, Alexey, et al.
Published: (2024)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
Temporal Inversion for Learning Interval Change in Chest X-Rays
by: Ko, Hanbin, et al.
Published: (2026)
by: Ko, Hanbin, et al.
Published: (2026)
DA-RAW: Domain Adaptive Object Detection for Real-World Adverse Weather Conditions
by: Jeon, Minsik, et al.
Published: (2023)
by: Jeon, Minsik, et al.
Published: (2023)
CXR-LT 2026 Challenge: Projection-Aware Multi-Label and Zero-Shot Chest X-Ray Classification
by: Cho, Juno, et al.
Published: (2026)
by: Cho, Juno, et al.
Published: (2026)
From Rays to Projections: Better Inputs for Feed-Forward View Synthesis
by: Wu, Zirui, et al.
Published: (2026)
by: Wu, Zirui, et al.
Published: (2026)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
UniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulation
by: Mittal, Himangi, et al.
Published: (2025)
by: Mittal, Himangi, et al.
Published: (2025)
Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
Novel View Synthesis as Video Completion
by: Wu, Qi, et al.
Published: (2026)
by: Wu, Qi, et al.
Published: (2026)
Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
by: Ling, Xingtao, et al.
Published: (2025)
by: Ling, Xingtao, et al.
Published: (2025)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
by: Tan, Jeff, et al.
Published: (2024)
by: Tan, Jeff, et al.
Published: (2024)
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
by: Vuong, Khiem, et al.
Published: (2025)
by: Vuong, Khiem, et al.
Published: (2025)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
RayEmb: Arbitrary Landmark Detection in X-Ray Images Using Ray Embedding Subspace
by: Shrestha, Pragyan, et al.
Published: (2024)
by: Shrestha, Pragyan, et al.
Published: (2024)
Similar Items
-
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026) -
LiTo: Surface Light Field Tokenization
by: Chang, Jen-Hao Rick, et al.
Published: (2026) -
Velox: Learning Representations of 4D Geometry and Appearance
by: Malik, Anagh, et al.
Published: (2026) -
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024) -
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)