Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Jian, Birsak, Michael, Cui, Wenqing, Li, Zhenyu, Wonka, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Any Resolution Any Geometry: From Multi-View To Multi-Patch
by: Cui, Wenqing, et al.
Published: (2026)
by: Cui, Wenqing, et al.
Published: (2026)
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
MatCLIP: Light- and Shape-Insensitive Assignment of PBR Material Models
by: Birsak, Michael, et al.
Published: (2025)
by: Birsak, Michael, et al.
Published: (2025)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
Patchwork: A compact representation for 3D polygonal shapes
by: Zheng, Ruichen, et al.
Published: (2026)
by: Zheng, Ruichen, et al.
Published: (2026)
Amodal Depth Anything: Amodal Depth Estimation in the Wild
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
efunc: An Efficient Function Representation without Neural Networks
by: Zhang, Biao, et al.
Published: (2025)
by: Zhang, Biao, et al.
Published: (2025)
Generative Human Geometry Distribution
by: Tang, Xiangjun, et al.
Published: (2025)
by: Tang, Xiangjun, et al.
Published: (2025)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
by: Choi, Jiho, et al.
Published: (2026)
by: Choi, Jiho, et al.
Published: (2026)
LASPA: Latent Spatial Alignment for Fast Training-free Single Image Editing
by: Alharbi, Yazeed, et al.
Published: (2024)
by: Alharbi, Yazeed, et al.
Published: (2024)
PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
by: Chen, Jianqi, et al.
Published: (2025)
by: Chen, Jianqi, et al.
Published: (2025)
PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
Geometry Distributions
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
by: Zeng, Yangchen, et al.
Published: (2026)
by: Zeng, Yangchen, et al.
Published: (2026)
LaGeM: A Large Geometry Model for 3D Representation Learning and Diffusion
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
Human Geometry Distribution for 3D Animation Generation
by: Tang, Xiangjun, et al.
Published: (2025)
by: Tang, Xiangjun, et al.
Published: (2025)
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Anatomical Positional Embeddings
by: Goncharov, Mikhail, et al.
Published: (2024)
by: Goncharov, Mikhail, et al.
Published: (2024)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
Shortcut Learning in Glomerular AI: Adversarial Penalties Hurt, Entropy Helps
by: Daouk, Mohammad, et al.
Published: (2026)
by: Daouk, Mohammad, et al.
Published: (2026)
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
by: Li, Yayuan, et al.
Published: (2026)
by: Li, Yayuan, et al.
Published: (2026)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
by: Lin, Yuhao, et al.
Published: (2024)
by: Lin, Yuhao, et al.
Published: (2024)
Positional Embedding-Aware Activations
by: Shah, Kathan, et al.
Published: (2023)
by: Shah, Kathan, et al.
Published: (2023)
Dissolving Is Amplifying: Towards Fine-Grained Anomaly Detection
by: Shi, Jian, et al.
Published: (2023)
by: Shi, Jian, et al.
Published: (2023)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
by: Eldesokey, Abdelrahman, et al.
Published: (2024)
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
by: Lu, Zhenyu, et al.
Published: (2025)
by: Lu, Zhenyu, et al.
Published: (2025)
PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD Reconstruction
by: Yang, Bingchen, et al.
Published: (2024)
by: Yang, Bingchen, et al.
Published: (2024)
Collaborative Position Reasoning Network for Referring Image Segmentation
by: Cao, Jianjian, et al.
Published: (2024)
by: Cao, Jianjian, et al.
Published: (2024)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
What Helps---and What Hurts: Bidirectional Explanations for Vision Transformers
by: Su, Qin, et al.
Published: (2026)
by: Su, Qin, et al.
Published: (2026)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
by: Hou, Jinghua, et al.
Published: (2024)
by: Hou, Jinghua, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
by: Cvejic, Aleksandar, et al.
Published: (2025)
by: Cvejic, Aleksandar, et al.
Published: (2025)
Similar Items
-
Any Resolution Any Geometry: From Multi-View To Multi-Patch
by: Cui, Wenqing, et al.
Published: (2026) -
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
by: Shi, Jian, et al.
Published: (2024) -
PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
by: Li, Zhenyu, et al.
Published: (2025) -
MatCLIP: Light- and Shape-Insensitive Assignment of PBR Material Models
by: Birsak, Michael, et al.
Published: (2025) -
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)