Continuous 3D Perception Model with Persistent State
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qianqian, Zhang, Yifei, Holynski, Aleksander, Efros, Alexei A., Kanazawa, Angjoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangled 3D Scene Generation with Layout Learning
by: Epstein, Dave, et al.
Published: (2024)
by: Epstein, Dave, et al.
Published: (2024)
Toon3D: Seeing Cartoons from New Perspectives
by: Weber, Ethan, et al.
Published: (2024)
by: Weber, Ethan, et al.
Published: (2024)
Rethinking Score Distillation as a Bridge Between Image Distributions
by: McAllister, David, et al.
Published: (2024)
by: McAllister, David, et al.
Published: (2024)
GPS as a Control Signal for Image Generation
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
Diffusion Models as Data Mining Tools
by: Siglidis, Ioannis, et al.
Published: (2024)
by: Siglidis, Ioannis, et al.
Published: (2024)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
by: Li, Zhengqi, et al.
Published: (2024)
by: Li, Zhengqi, et al.
Published: (2024)
Self-Improving 4D Perception via Self-Distillation
by: Huang, Nan, et al.
Published: (2026)
by: Huang, Nan, et al.
Published: (2026)
Generating Continual Human Motion in Diverse 3D Scenes
by: Mir, Aymen, et al.
Published: (2023)
by: Mir, Aymen, et al.
Published: (2023)
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026)
by: Bonnen, Tyler, et al.
Published: (2026)
Shape of Motion: 4D Reconstruction from a Single Video
by: Wang, Qianqian, et al.
Published: (2024)
by: Wang, Qianqian, et al.
Published: (2024)
The More You See in 2D, the More You Perceive in 3D
by: Han, Xinyang, et al.
Published: (2024)
by: Han, Xinyang, et al.
Published: (2024)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
by: Xu, Congrong, et al.
Published: (2024)
by: Xu, Congrong, et al.
Published: (2024)
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
by: Pan, Zhuoyang, et al.
Published: (2024)
by: Pan, Zhuoyang, et al.
Published: (2024)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
by: Kerr, Justin, et al.
Published: (2024)
by: Kerr, Justin, et al.
Published: (2024)
COLMAP-Free 3D Gaussian Splatting
by: Fu, Yang, et al.
Published: (2023)
by: Fu, Yang, et al.
Published: (2023)
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
by: Jin, Linyi, et al.
Published: (2024)
by: Jin, Linyi, et al.
Published: (2024)
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting
by: Bhattad, Anand, et al.
Published: (2025)
by: Bhattad, Anand, et al.
Published: (2025)
Interpreting the Second-Order Effects of Neurons in CLIP
by: Gandelsman, Yossi, et al.
Published: (2024)
by: Gandelsman, Yossi, et al.
Published: (2024)
CAT3D: Create Anything in 3D with Multi-View Diffusion Models
by: Gao, Ruiqi, et al.
Published: (2024)
by: Gao, Ruiqi, et al.
Published: (2024)
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
by: Yang, Gengshan, et al.
Published: (2024)
by: Yang, Gengshan, et al.
Published: (2024)
Generative Image Dynamics
by: Li, Zhengqi, et al.
Published: (2023)
by: Li, Zhengqi, et al.
Published: (2023)
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
by: Wu, Rundi, et al.
Published: (2024)
by: Wu, Rundi, et al.
Published: (2024)
Infinite Texture: Text-guided High Resolution Diffusion Texture Synthesis
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Bolt3D: Generating 3D Scenes in Seconds
by: Szymanowicz, Stanislaw, et al.
Published: (2025)
by: Szymanowicz, Stanislaw, et al.
Published: (2025)
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
by: Jin, Haian, et al.
Published: (2026)
by: Jin, Haian, et al.
Published: (2026)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
Video Interpolation with Diffusion Models
by: Jain, Siddhant, et al.
Published: (2024)
by: Jain, Siddhant, et al.
Published: (2024)
Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
by: Wu, Mingxuan, et al.
Published: (2025)
by: Wu, Mingxuan, et al.
Published: (2025)
Readout Guidance: Learning Control from Diffusion Features
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
Synthesizing Moving People with 3D Control
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024)
by: Müller, Lea, et al.
Published: (2024)
How Animals Dance (When You're Not Looking)
by: Wang, Xiaojuan, et al.
Published: (2025)
by: Wang, Xiaojuan, et al.
Published: (2025)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
by: Wang, Xiaojuan, et al.
Published: (2024)
by: Wang, Xiaojuan, et al.
Published: (2024)
IT$^3$: Idempotent Test-Time Training
by: Durasov, Nikita, et al.
Published: (2024)
by: Durasov, Nikita, et al.
Published: (2024)
Decentralized Diffusion Models
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
Similar Items
-
Disentangled 3D Scene Generation with Layout Learning
by: Epstein, Dave, et al.
Published: (2024) -
Toon3D: Seeing Cartoons from New Perspectives
by: Weber, Ethan, et al.
Published: (2024) -
Rethinking Score Distillation as a Bridge Between Image Distributions
by: McAllister, David, et al.
Published: (2024) -
GPS as a Control Signal for Image Generation
by: Feng, Chao, et al.
Published: (2025) -
Diffusion Models as Data Mining Tools
by: Siglidis, Ioannis, et al.
Published: (2024)