St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
Fuente:
arXiv
Guardado en:
| Autores principales: | Feng, Haiwen, Zhang, Junyi, Wang, Qianqian, Ye, Yufei, Yu, Pengcheng, Black, Michael J., Darrell, Trevor, Kanazawa, Angjoo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-Improving 4D Perception via Self-Distillation
por: Huang, Nan, et al.
Publicado: (2026)
por: Huang, Nan, et al.
Publicado: (2026)
Shape of Motion: 4D Reconstruction from a Single Video
por: Wang, Qianqian, et al.
Publicado: (2024)
por: Wang, Qianqian, et al.
Publicado: (2024)
Predicting 4D Hand Trajectory from Monocular Videos
por: Ye, Yufei, et al.
Publicado: (2025)
por: Ye, Yufei, et al.
Publicado: (2025)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
por: Yin, Shaofeng, et al.
Publicado: (2026)
por: Yin, Shaofeng, et al.
Publicado: (2026)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
por: Kerr, Justin, et al.
Publicado: (2024)
por: Kerr, Justin, et al.
Publicado: (2024)
Visually Prompted Benchmarks Are Surprisingly Fragile
por: Feng, Haiwen, et al.
Publicado: (2025)
por: Feng, Haiwen, et al.
Publicado: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
por: Ng, Evonne, et al.
Publicado: (2024)
por: Ng, Evonne, et al.
Publicado: (2024)
Continuous 3D Perception Model with Persistent State
por: Wang, Qianqian, et al.
Publicado: (2025)
por: Wang, Qianqian, et al.
Publicado: (2025)
Human-level 3D shape perception emerges from multi-view learning
por: Bonnen, Tyler, et al.
Publicado: (2026)
por: Bonnen, Tyler, et al.
Publicado: (2026)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
por: Xu, Congrong, et al.
Publicado: (2024)
por: Xu, Congrong, et al.
Publicado: (2024)
Visual Imitation Enables Contextual Humanoid Control
por: Allshire, Arthur, et al.
Publicado: (2025)
por: Allshire, Arthur, et al.
Publicado: (2025)
Reconstructing People, Places, and Cameras
por: Müller, Lea, et al.
Publicado: (2024)
por: Müller, Lea, et al.
Publicado: (2024)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
por: Pan, Zhuoyang, et al.
Publicado: (2024)
por: Pan, Zhuoyang, et al.
Publicado: (2024)
Generating Continual Human Motion in Diverse 3D Scenes
por: Mir, Aymen, et al.
Publicado: (2023)
por: Mir, Aymen, et al.
Publicado: (2023)
The More You See in 2D, the More You Perceive in 3D
por: Han, Xinyang, et al.
Publicado: (2024)
por: Han, Xinyang, et al.
Publicado: (2024)
Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
por: Wu, Mingxuan, et al.
Publicado: (2025)
por: Wu, Mingxuan, et al.
Publicado: (2025)
Segment Any Motion in Videos
por: Huang, Nan, et al.
Publicado: (2025)
por: Huang, Nan, et al.
Publicado: (2025)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
por: Li, Zhengqi, et al.
Publicado: (2024)
por: Li, Zhengqi, et al.
Publicado: (2024)
Toward Human Understanding with Controllable Synthesis
por: Cuevas-Velasquez, Hanz, et al.
Publicado: (2024)
por: Cuevas-Velasquez, Hanz, et al.
Publicado: (2024)
WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion
por: Shin, Soyong, et al.
Publicado: (2023)
por: Shin, Soyong, et al.
Publicado: (2023)
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
por: Yang, Gengshan, et al.
Publicado: (2024)
por: Yang, Gengshan, et al.
Publicado: (2024)
Toon3D: Seeing Cartoons from New Perspectives
por: Weber, Ethan, et al.
Publicado: (2024)
por: Weber, Ethan, et al.
Publicado: (2024)
Estimating Body and Hand Motion in an Ego-sensed World
por: Yi, Brent, et al.
Publicado: (2024)
por: Yi, Brent, et al.
Publicado: (2024)
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
por: Jiang, Zeren, et al.
Publicado: (2026)
por: Jiang, Zeren, et al.
Publicado: (2026)
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
por: Qian, Shenhan, et al.
Publicado: (2026)
por: Qian, Shenhan, et al.
Publicado: (2026)
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
por: Lu, Jiahao, et al.
Publicado: (2026)
por: Lu, Jiahao, et al.
Publicado: (2026)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
por: Plizzari, Chiara, et al.
Publicado: (2024)
por: Plizzari, Chiara, et al.
Publicado: (2024)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
por: Akkerman, Rick, et al.
Publicado: (2024)
por: Akkerman, Rick, et al.
Publicado: (2024)
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
por: Zhang, Junyi, et al.
Publicado: (2026)
por: Zhang, Junyi, et al.
Publicado: (2026)
NeRF-XL: Scaling NeRFs with Multiple GPUs
por: Li, Ruilong, et al.
Publicado: (2024)
por: Li, Ruilong, et al.
Publicado: (2024)
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
por: Athanasiou, Nikos, et al.
Publicado: (2023)
por: Athanasiou, Nikos, et al.
Publicado: (2023)
Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
por: Wu, Junyi, et al.
Publicado: (2025)
por: Wu, Junyi, et al.
Publicado: (2025)
Fillerbuster: Unified Generative Scene Completion Model for Casual Captures
por: Weber, Ethan, et al.
Publicado: (2025)
por: Weber, Ethan, et al.
Publicado: (2025)
GenLit: Reformulating Single-Image Relighting as Video Generation
por: Bharadwaj, Shrisha, et al.
Publicado: (2024)
por: Bharadwaj, Shrisha, et al.
Publicado: (2024)
Re-Thinking Inverse Graphics With Large Language Models
por: Kulits, Peter, et al.
Publicado: (2024)
por: Kulits, Peter, et al.
Publicado: (2024)
Synergy and Synchrony in Couple Dances
por: Maluleke, Vongani, et al.
Publicado: (2024)
por: Maluleke, Vongani, et al.
Publicado: (2024)
Laplacian Analysis Meets Dynamics Modelling: Gaussian Splatting for 4D Reconstruction
por: Zhou, Yifan, et al.
Publicado: (2025)
por: Zhou, Yifan, et al.
Publicado: (2025)
UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video
por: Sur, Tanuj, et al.
Publicado: (2026)
por: Sur, Tanuj, et al.
Publicado: (2026)
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
por: Tang, George, et al.
Publicado: (2025)
por: Tang, George, et al.
Publicado: (2025)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
por: Maluleke, Vongani H., et al.
Publicado: (2025)
por: Maluleke, Vongani H., et al.
Publicado: (2025)
Ejemplares similares
-
Self-Improving 4D Perception via Self-Distillation
por: Huang, Nan, et al.
Publicado: (2026) -
Shape of Motion: 4D Reconstruction from a Single Video
por: Wang, Qianqian, et al.
Publicado: (2024) -
Predicting 4D Hand Trajectory from Monocular Videos
por: Ye, Yufei, et al.
Publicado: (2025) -
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
por: Yin, Shaofeng, et al.
Publicado: (2026) -
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
por: Kerr, Justin, et al.
Publicado: (2024)