MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Junyi, Herrmann, Charles, Hur, Junhwa, Jampani, Varun, Darrell, Trevor, Cole, Forrester, Sun, Deqing, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
by: Zhang, Junyi, et al.
Published: (2026)
by: Zhang, Junyi, et al.
Published: (2026)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023)
by: Zhang, Junyi, et al.
Published: (2023)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025)
by: Gu, Leslie, et al.
Published: (2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories
by: Geng, Daniel, et al.
Published: (2024)
by: Geng, Daniel, et al.
Published: (2024)
Boundary Attention: Learning curves, corners, junctions and grouping
by: Polansky, Mia Gaia, et al.
Published: (2024)
by: Polansky, Mia Gaia, et al.
Published: (2024)
WonderJourney: Going from Anywhere to Everywhere
by: Yu, Hong-Xing, et al.
Published: (2023)
by: Yu, Hong-Xing, et al.
Published: (2023)
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024)
by: Zhong, Lei, et al.
Published: (2024)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
by: Hur, Junhwa, et al.
Published: (2026)
by: Hur, Junhwa, et al.
Published: (2026)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
A Simple Approach to Unifying Diffusion-based Conditional Generation
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
MotionV2V: Editing Motion in a Video
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
by: Shah, Viraj, et al.
Published: (2023)
by: Shah, Viraj, et al.
Published: (2023)
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
by: Hur, Junhwa, et al.
Published: (2024)
by: Hur, Junhwa, et al.
Published: (2024)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
MASIV: Toward Material-Agnostic System Identification from Videos
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
DreamWalk: Style Space Exploration using Diffusion Guidance
by: Shu, Michelle, et al.
Published: (2024)
by: Shu, Michelle, et al.
Published: (2024)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
MonSTeR: a Unified Model for Motion, Scene, Text Retrieval
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
by: Nam, Jisu, et al.
Published: (2026)
by: Nam, Jisu, et al.
Published: (2026)
Human Video Generation from a Single Image with 3D Pose and View Control
by: Wang, Tiantian, et al.
Published: (2026)
by: Wang, Tiantian, et al.
Published: (2026)
Global Diffusive Expansion of Boltzmann Equation in exterior Domain
by: Jung, Junhwa
Published: (2023)
by: Jung, Junhwa
Published: (2023)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement
by: Boss, Mark, et al.
Published: (2024)
by: Boss, Mark, et al.
Published: (2024)
ReSWD: ReSTIR'd, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction
by: Boss, Mark, et al.
Published: (2025)
by: Boss, Mark, et al.
Published: (2025)
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
by: Xiong, Yajiao, et al.
Published: (2025)
by: Xiong, Yajiao, et al.
Published: (2025)
Four Simple Proprioceptive Estimators for Legged Robots
by: Dellaert, Frank, et al.
Published: (2026)
by: Dellaert, Frank, et al.
Published: (2026)
Repulsive Trajectory Modification and Conflict Resolution for Efficient Multi-Manipulator Motion Planning
by: Hong, Junhwa, et al.
Published: (2025)
by: Hong, Junhwa, et al.
Published: (2025)
MARBLE: Material Recomposition and Blending in CLIP-Space
by: Cheng, Ta-Ying, et al.
Published: (2025)
by: Cheng, Ta-Ying, et al.
Published: (2025)
FROMAT: Multiview Material Appearance Transfer via Few-Shot Self-Attention Adaptation
by: Kompanowski, Hubert, et al.
Published: (2025)
by: Kompanowski, Hubert, et al.
Published: (2025)
Lumiere: A Space-Time Diffusion Model for Video Generation
by: Bar-Tal, Omer, et al.
Published: (2024)
by: Bar-Tal, Omer, et al.
Published: (2024)
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
by: Gillman, Nate, et al.
Published: (2025)
by: Gillman, Nate, et al.
Published: (2025)
Probing the 3D Awareness of Visual Foundation Models
by: Banani, Mohamed El, et al.
Published: (2024)
by: Banani, Mohamed El, et al.
Published: (2024)
Self-training Room Layout Estimation via Geometry-aware Ray-casting
by: Solarte, Bolivar, et al.
Published: (2024)
by: Solarte, Bolivar, et al.
Published: (2024)
Topological Matter and Fractional Entangled Quantum Geometry through Light
by: Hur, Karyn Le
Published: (2022)
by: Hur, Karyn Le
Published: (2022)
Simple Hardware Implementation of Motion Estimation Algorithms
by: Juan Romero
Published: (2019)
by: Juan Romero
Published: (2019)
Similar Items
-
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
by: Zhang, Junyi, et al.
Published: (2026) -
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023) -
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025) -
Motion Prompting: Controlling Video Generation with Motion Trajectories
by: Geng, Daniel, et al.
Published: (2024) -
Boundary Attention: Learning curves, corners, junctions and grouping
by: Polansky, Mia Gaia, et al.
Published: (2024)