DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qitao, Lin, Amy, Tan, Jeff, Zhang, Jason Y., Ramanan, Deva, Tulsiani, Shubham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cameras as Rays: Pose Estimation via Ray Diffusion
by: Zhang, Jason Y., et al.
Published: (2024)
by: Zhang, Jason Y., et al.
Published: (2024)
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024)
by: Zhao, Qitao, et al.
Published: (2024)
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
by: Tan, Jeff, et al.
Published: (2024)
by: Tan, Jeff, et al.
Published: (2024)
CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026)
by: Cong, Zhongxiao, et al.
Published: (2026)
AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
by: Vuong, Khiem, et al.
Published: (2025)
by: Vuong, Khiem, et al.
Published: (2025)
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
Using Diffusion Priors for Video Amodal Segmentation
by: Chen, Kaihua, et al.
Published: (2024)
by: Chen, Kaihua, et al.
Published: (2024)
Predicting Long-horizon Futures by Conditioning on Geometry and Time
by: Khurana, Tarasha, et al.
Published: (2024)
by: Khurana, Tarasha, et al.
Published: (2024)
MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
LightSwitch: Multi-view Relighting with Material-guided Diffusion
by: Litman, Yehonathan, et al.
Published: (2025)
by: Litman, Yehonathan, et al.
Published: (2025)
SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
by: Bokhovkin, Alexey, et al.
Published: (2024)
by: Bokhovkin, Alexey, et al.
Published: (2024)
CuSfM: CUDA-Accelerated Structure-from-Motion
by: Yu, Jingrui, et al.
Published: (2025)
by: Yu, Jingrui, et al.
Published: (2025)
Dense-SfM: Structure from Motion with Dense Consistent Matching
by: Lee, JongMin, et al.
Published: (2025)
by: Lee, JongMin, et al.
Published: (2025)
RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
by: Duisterhof, Bardienus P., et al.
Published: (2025)
by: Duisterhof, Bardienus P., et al.
Published: (2025)
UpFusion: Novel View Diffusion from Unposed Sparse View Observations
by: Kani, Bharath Raj Nagoor, et al.
Published: (2023)
by: Kani, Bharath Raj Nagoor, et al.
Published: (2023)
ColabSfM: Collaborative Structure-from-Motion by Point Cloud Registration
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
Accenture-NVS1: A Novel View Synthesis Dataset
by: Sugg, Thomas, et al.
Published: (2025)
by: Sugg, Thomas, et al.
Published: (2025)
MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
by: Pataki, Zador, et al.
Published: (2025)
by: Pataki, Zador, et al.
Published: (2025)
DP-SfM: Dual-Pixel Structure-from-Motion without Scale Ambiguity
by: Makabe, Lilika, et al.
Published: (2026)
by: Makabe, Lilika, et al.
Published: (2026)
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
by: Wu, Yu, et al.
Published: (2026)
by: Wu, Yu, et al.
Published: (2026)
Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
by: Chen, Kaihua, et al.
Published: (2025)
by: Chen, Kaihua, et al.
Published: (2025)
MaterialFusion: Enhancing Inverse Rendering with Material Diffusion Priors
by: Litman, Yehonathan, et al.
Published: (2024)
by: Litman, Yehonathan, et al.
Published: (2024)
Light3R-SfM: Towards Feed-forward Structure-from-Motion
by: Elflein, Sven, et al.
Published: (2025)
by: Elflein, Sven, et al.
Published: (2025)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
DATAP-SfM: Dynamic-Aware Tracking Any Point for Robust Structure from Motion in the Wild
by: Ye, Weicai, et al.
Published: (2024)
by: Ye, Weicai, et al.
Published: (2024)
MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
by: Duisterhof, Bardienus, et al.
Published: (2024)
by: Duisterhof, Bardienus, et al.
Published: (2024)
LiVisSfM: Accurate and Robust Structure-from-Motion with LiDAR and Visual Cues
by: Jiang, Hanqing, et al.
Published: (2024)
by: Jiang, Hanqing, et al.
Published: (2024)
Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment
by: Gao, Yurong, et al.
Published: (2026)
by: Gao, Yurong, et al.
Published: (2026)
InstantSfM: Towards GPU-Native SfM for the Deep Learning Era
by: Zhong, Jiankun, et al.
Published: (2025)
by: Zhong, Jiankun, et al.
Published: (2025)
RefAV: Towards Planning-Centric Scenario Mining
by: Davidson, Cainan, et al.
Published: (2025)
by: Davidson, Cainan, et al.
Published: (2025)
Revisiting Few-Shot Object Detection with Vision-Language Models
by: Madan, Anish, et al.
Published: (2023)
by: Madan, Anish, et al.
Published: (2023)
SMORE: Simultaneous Map and Object REconstruction
by: Chodosh, Nathaniel, et al.
Published: (2024)
by: Chodosh, Nathaniel, et al.
Published: (2024)
Tackling the Singularities at the Endpoints of Time Intervals in Diffusion Models
by: Zhang, Pengze, et al.
Published: (2024)
by: Zhang, Pengze, et al.
Published: (2024)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
Diverse Score Distillation
by: Xu, Yanbo, et al.
Published: (2024)
by: Xu, Yanbo, et al.
Published: (2024)
Generating Physically Stable and Buildable Brick Structures from Text
by: Pun, Ava, et al.
Published: (2025)
by: Pun, Ava, et al.
Published: (2025)
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
by: Deng, Junyuan, et al.
Published: (2025)
by: Deng, Junyuan, et al.
Published: (2025)
Similar Items
-
Cameras as Rays: Pose Estimation via Ray Diffusion
by: Zhang, Jason Y., et al.
Published: (2024) -
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
by: Zhao, Qitao, et al.
Published: (2024) -
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
by: Tan, Jeff, et al.
Published: (2024) -
CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
by: Wang, Zihan, et al.
Published: (2025) -
Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning
by: Cong, Zhongxiao, et al.
Published: (2026)