Human-level 3D shape perception emerges from multi-view learning
Fuente:
arXiv
Saved in:
| Main Authors: | Bonnen, Tyler, Malik, Jitendra, Kanazawa, Angjoo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024)
by: Müller, Lea, et al.
Published: (2024)
Approaching human 3D shape perception with neurally mappable models
by: O'Connell, Thomas P., et al.
Published: (2023)
by: O'Connell, Thomas P., et al.
Published: (2023)
Generating Continual Human Motion in Diverse 3D Scenes
by: Mir, Aymen, et al.
Published: (2023)
by: Mir, Aymen, et al.
Published: (2023)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)
by: Maluleke, Vongani, et al.
Published: (2024)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
by: Maluleke, Vongani H., et al.
Published: (2025)
by: Maluleke, Vongani H., et al.
Published: (2025)
The More You See in 2D, the More You Perceive in 3D
by: Han, Xinyang, et al.
Published: (2024)
by: Han, Xinyang, et al.
Published: (2024)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
by: Xu, Congrong, et al.
Published: (2024)
by: Xu, Congrong, et al.
Published: (2024)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
by: Pan, Zhuoyang, et al.
Published: (2024)
by: Pan, Zhuoyang, et al.
Published: (2024)
Continuous 3D Perception Model with Persistent State
by: Wang, Qianqian, et al.
Published: (2025)
by: Wang, Qianqian, et al.
Published: (2025)
Toon3D: Seeing Cartoons from New Perspectives
by: Weber, Ethan, et al.
Published: (2024)
by: Weber, Ethan, et al.
Published: (2024)
Estimating Body and Hand Motion in an Ego-sensed World
by: Yi, Brent, et al.
Published: (2024)
by: Yi, Brent, et al.
Published: (2024)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
by: Ng, Evonne, et al.
Published: (2024)
by: Ng, Evonne, et al.
Published: (2024)
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
by: Yang, Gengshan, et al.
Published: (2024)
by: Yang, Gengshan, et al.
Published: (2024)
Reconstructing Hand-Held Objects in 3D from Images and Videos
by: Wu, Jane, et al.
Published: (2024)
by: Wu, Jane, et al.
Published: (2024)
Shape of Motion: 4D Reconstruction from a Single Video
by: Wang, Qianqian, et al.
Published: (2024)
by: Wang, Qianqian, et al.
Published: (2024)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
Tracking by Predicting 3-D Gaussians Over Time
by: Baranwal, Tanish, et al.
Published: (2025)
by: Baranwal, Tanish, et al.
Published: (2025)
Adaptive Human Trajectory Prediction via Latent Corridors
by: Thakkar, Neerja, et al.
Published: (2023)
by: Thakkar, Neerja, et al.
Published: (2023)
Self-Improving 4D Perception via Self-Distillation
by: Huang, Nan, et al.
Published: (2026)
by: Huang, Nan, et al.
Published: (2026)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Evaluating Multiview Object Consistency in Humans and Image Models
by: Bonnen, Tyler, et al.
Published: (2024)
by: Bonnen, Tyler, et al.
Published: (2024)
Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
by: Wu, Mingxuan, et al.
Published: (2025)
by: Wu, Mingxuan, et al.
Published: (2025)
NeRF-XL: Scaling NeRFs with Multiple GPUs
by: Li, Ruilong, et al.
Published: (2024)
by: Li, Ruilong, et al.
Published: (2024)
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Self-learning Canonical Space for Multi-view 3D Human Pose Estimation
by: Li, Xiaoben, et al.
Published: (2024)
by: Li, Xiaoben, et al.
Published: (2024)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
by: Kerr, Justin, et al.
Published: (2024)
by: Kerr, Justin, et al.
Published: (2024)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
Efficient multi-view training for 3D Gaussian Splatting
by: Choi, Minhyuk, et al.
Published: (2025)
by: Choi, Minhyuk, et al.
Published: (2025)
Decentralized Diffusion Models
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
GARField: Group Anything with Radiance Fields
by: Kim, Chung Min, et al.
Published: (2024)
by: Kim, Chung Min, et al.
Published: (2024)
Cameras as Relative Positional Encoding
by: Li, Ruilong, et al.
Published: (2025)
by: Li, Ruilong, et al.
Published: (2025)
Viser: Imperative, Web-based 3D Visualization in Python
by: Yi, Brent, et al.
Published: (2025)
by: Yi, Brent, et al.
Published: (2025)
MV-MR: multi-views and multi-representations for self-supervised learning and knowledge distillation
by: Kinakh, Vitaliy, et al.
Published: (2023)
by: Kinakh, Vitaliy, et al.
Published: (2023)
Fillerbuster: Unified Generative Scene Completion Model for Casual Captures
by: Weber, Ethan, et al.
Published: (2025)
by: Weber, Ethan, et al.
Published: (2025)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
by: Li, Zhengqi, et al.
Published: (2024)
by: Li, Zhengqi, et al.
Published: (2024)
Multi-view learning for automatic classification of multi-wavelength auroral images
by: Yang, Qiuju, et al.
Published: (2023)
by: Yang, Qiuju, et al.
Published: (2023)
SAM 3D Body: Robust Full-Body Human Mesh Recovery
by: Yang, Xitong, et al.
Published: (2026)
by: Yang, Xitong, et al.
Published: (2026)
MPL: Lifting 3D Human Pose from Multi-view 2D Poses
by: Ghasemzadeh, Seyed Abolfazl, et al.
Published: (2024)
by: Ghasemzadeh, Seyed Abolfazl, et al.
Published: (2024)
Similar Items
-
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024) -
Approaching human 3D shape perception with neurally mappable models
by: O'Connell, Thomas P., et al.
Published: (2023) -
Generating Continual Human Motion in Diverse 3D Scenes
by: Mir, Aymen, et al.
Published: (2023) -
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025) -
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)