ViPE: Video Pose Engine for 3D Geometric Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jiahui, Zhou, Qunjie, Rabeti, Hesam, Korovko, Aleksandr, Ling, Huan, Ren, Xuanchi, Shen, Tianchang, Gao, Jun, Slepichev, Dmitry, Lin, Chen-Hsuan, Ren, Jiawei, Xie, Kevin, Biswas, Joydeep, Leal-Taixe, Laura, Fidler, Sanja |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
cuVSLAM: CUDA accelerated visual odometry and mapping
by: Korovko, Alexander, et al.
Published: (2025)
by: Korovko, Alexander, et al.
Published: (2025)
Depth Completion as Parameter-Efficient Test-Time Adaptation
by: Ke, Bingxin, et al.
Published: (2026)
by: Ke, Bingxin, et al.
Published: (2026)
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026)
by: Liu, Shaowei, et al.
Published: (2026)
XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
by: Ren, Xuanchi, et al.
Published: (2023)
by: Ren, Xuanchi, et al.
Published: (2023)
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
by: Lu, Yifan, et al.
Published: (2024)
by: Lu, Yifan, et al.
Published: (2024)
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
by: Bahmani, Sherwin, et al.
Published: (2025)
by: Bahmani, Sherwin, et al.
Published: (2025)
Industrial cuVSLAM Benchmark & Integration
by: Hana, Charbel Abi, et al.
Published: (2026)
by: Hana, Charbel Abi, et al.
Published: (2026)
SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
by: Ren, Xuanchi, et al.
Published: (2024)
by: Ren, Xuanchi, et al.
Published: (2024)
Lyra 2.0: Explorable Generative 3D Worlds
by: Shen, Tianchang, et al.
Published: (2026)
by: Shen, Tianchang, et al.
Published: (2026)
The NeRFect Match: Exploring NeRF Features for Visual Localization
by: Zhou, Qunjie, et al.
Published: (2024)
by: Zhou, Qunjie, et al.
Published: (2024)
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
by: Lu, Yifan, et al.
Published: (2026)
by: Lu, Yifan, et al.
Published: (2026)
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
by: Nasser, Zaid, et al.
Published: (2025)
by: Nasser, Zaid, et al.
Published: (2025)
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
by: Ren, Xuanchi, et al.
Published: (2025)
by: Ren, Xuanchi, et al.
Published: (2025)
MATCHA:Towards Matching Anything
by: Xue, Fei, et al.
Published: (2025)
by: Xue, Fei, et al.
Published: (2025)
Light3R-SfM: Towards Feed-forward Structure-from-Motion
by: Elflein, Sven, et al.
Published: (2025)
by: Elflein, Sven, et al.
Published: (2025)
RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments
by: Nasser, Zaid, et al.
Published: (2026)
by: Nasser, Zaid, et al.
Published: (2026)
APE: Agentic Prompt Enhancer for Image Generation and Editing
by: Huang, Zijian, et al.
Published: (2026)
by: Huang, Zijian, et al.
Published: (2026)
DynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction
by: Seidenschwarz, Jenny, et al.
Published: (2024)
by: Seidenschwarz, Jenny, et al.
Published: (2024)
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
by: Liao, Yuan-Hong, et al.
Published: (2025)
by: Liao, Yuan-Hong, et al.
Published: (2025)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
Motion Attribution for Video Generation
by: Wu, Xindi, et al.
Published: (2026)
by: Wu, Xindi, et al.
Published: (2026)
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
by: Burzio, Alessandro, et al.
Published: (2026)
by: Burzio, Alessandro, et al.
Published: (2026)
Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
by: Liang, Hanxue, et al.
Published: (2024)
by: Liang, Hanxue, et al.
Published: (2024)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
fVDB: A Deep-Learning Framework for Sparse, Large-Scale, and High-Performance Spatial Intelligence
by: Williams, Francis, et al.
Published: (2024)
by: Williams, Francis, et al.
Published: (2024)
A Guide to Structureless Visual Localization
by: Panek, Vojtech, et al.
Published: (2025)
by: Panek, Vojtech, et al.
Published: (2025)
VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale
by: Elflein, Sven, et al.
Published: (2026)
by: Elflein, Sven, et al.
Published: (2026)
SYNAPSE: SYmbolic Neural-Aided Preference Synthesis Engine
by: Modak, Sadanand, et al.
Published: (2024)
by: Modak, Sadanand, et al.
Published: (2024)
SpaceMesh: A Continuous Representation for Learning Manifold Surface Meshes
by: Shen, Tianchang, et al.
Published: (2024)
by: Shen, Tianchang, et al.
Published: (2024)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
Align Your Steps: Optimizing Sampling Schedules in Diffusion Models
by: Sabour, Amirmojtaba, et al.
Published: (2024)
by: Sabour, Amirmojtaba, et al.
Published: (2024)
Manipulating the metal-insulator transitions in correlated vanadium dioxide through bandwidth and band-filling control
by: Yao, Xiaohui, et al.
Published: (2025)
by: Yao, Xiaohui, et al.
Published: (2025)
Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
by: Ling, Huan, et al.
Published: (2023)
by: Ling, Huan, et al.
Published: (2023)
Trajeglish: Traffic Modeling as Next-Token Prediction
by: Philion, Jonah, et al.
Published: (2023)
by: Philion, Jonah, et al.
Published: (2023)
L4GM: Large 4D Gaussian Reconstruction Model
by: Ren, Jiawei, et al.
Published: (2024)
by: Ren, Jiawei, et al.
Published: (2024)
Native Segmentation Vision Transformers
by: Brasó, Guillem, et al.
Published: (2025)
by: Brasó, Guillem, et al.
Published: (2025)
Univariate Bicycle Quantum LDPC Codes: Explicit Logical Structure and Distance Bounds
by: Rabeti, Sheida, et al.
Published: (2026)
by: Rabeti, Sheida, et al.
Published: (2026)
Similar Items
-
cuVSLAM: CUDA accelerated visual odometry and mapping
by: Korovko, Alexander, et al.
Published: (2025) -
Depth Completion as Parameter-Efficient Test-Time Adaptation
by: Ke, Bingxin, et al.
Published: (2026) -
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026) -
XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
by: Ren, Xuanchi, et al.
Published: (2023) -
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
by: Ren, Xuanchi, et al.
Published: (2025)