Reconstructing Hand-Held Objects in 3D from Images and Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jane, Pavlakos, Georgios, Gkioxari, Georgia, Malik, Jitendra |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Real3D: Scaling Up Large Reconstruction Models with Real-World Images
by: Jiang, Hanwen, et al.
Published: (2024)
by: Jiang, Hanwen, et al.
Published: (2024)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025)
by: Karvounas, Giorgos, et al.
Published: (2025)
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025)
by: Sahoo, Aadarsh, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
by: Sahoo, Aadarsh, et al.
Published: (2026)
by: Sahoo, Aadarsh, et al.
Published: (2026)
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024)
by: Ma, Ziqi, et al.
Published: (2024)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
by: Fu, Bowen, et al.
Published: (2023)
by: Fu, Bowen, et al.
Published: (2023)
Estimating Body and Hand Motion in an Ego-sensed World
by: Yi, Brent, et al.
Published: (2024)
by: Yi, Brent, et al.
Published: (2024)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)
by: Maluleke, Vongani, et al.
Published: (2024)
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Hand Held Multi-Object Tracking Dataset in American Football
by: Otsubo, Rintaro, et al.
Published: (2025)
by: Otsubo, Rintaro, et al.
Published: (2025)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
by: Demler, Ilona, et al.
Published: (2025)
by: Demler, Ilona, et al.
Published: (2025)
ExpertAF: Expert Actionable Feedback from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Reconstructing Humans with a Biomechanically Accurate Skeleton
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
by: Chen, Yuantao, et al.
Published: (2026)
by: Chen, Yuantao, et al.
Published: (2026)
Depth Restoration of Hand-Held Transparent Objects for Human-to-Robot Handover
by: Yu, Ran, et al.
Published: (2024)
by: Yu, Ran, et al.
Published: (2024)
GHOST: Fast Category-agnostic Hand-Object Interaction Reconstruction from RGB Videos using Gaussian Splatting
by: Aboukhadra, Ahmed Tawfik, et al.
Published: (2026)
by: Aboukhadra, Ahmed Tawfik, et al.
Published: (2026)
Expressive Gaussian Human Avatars from Monocular RGB Video
by: Hu, Hezhen, et al.
Published: (2024)
by: Hu, Hezhen, et al.
Published: (2024)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
Atlas Gaussians Diffusion for 3D Generation
by: Yang, Haitao, et al.
Published: (2024)
by: Yang, Haitao, et al.
Published: (2024)
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026)
by: Bonnen, Tyler, et al.
Published: (2026)
HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions
by: Xu, Hao, et al.
Published: (2024)
by: Xu, Hao, et al.
Published: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026)
by: Kang, Raphi, et al.
Published: (2026)
Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos
by: Wei, Dingkun, et al.
Published: (2026)
by: Wei, Dingkun, et al.
Published: (2026)
HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
by: Wang, Shuaibing, et al.
Published: (2024)
by: Wang, Shuaibing, et al.
Published: (2024)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
TexHOI: Reconstructing Textures of 3D Unknown Objects in Monocular Hand-Object Interaction Scenes
by: Aggarwal, Alakh, et al.
Published: (2025)
by: Aggarwal, Alakh, et al.
Published: (2025)
Get a Grip: Reconstructing Hand-Object Stable Grasps in Egocentric Videos
by: Zhu, Zhifan, et al.
Published: (2023)
by: Zhu, Zhifan, et al.
Published: (2023)
HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
by: Hu, Hezhen, et al.
Published: (2026)
by: Hu, Hezhen, et al.
Published: (2026)
HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image
by: Choi, Hongsuk, et al.
Published: (2023)
by: Choi, Hongsuk, et al.
Published: (2023)
ShapeGraFormer: GraFormer-Based Network for Hand-Object Reconstruction from a Single Depth Map
by: Aboukhadra, Ahmed Tawfik, et al.
Published: (2023)
by: Aboukhadra, Ahmed Tawfik, et al.
Published: (2023)
CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction
by: Xu, Hao, et al.
Published: (2026)
by: Xu, Hao, et al.
Published: (2026)
Is CLIP ideal? No. Can we fix it? Yes!
by: Kang, Raphi, et al.
Published: (2025)
by: Kang, Raphi, et al.
Published: (2025)
Similar Items
-
Real3D: Scaling Up Large Reconstruction Models with Real-World Images
by: Jiang, Hanwen, et al.
Published: (2024) -
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025) -
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025) -
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024) -
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)