HORT: Monocular Hand-held Objects Reconstruction with Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zerui, Potamias, Rolandos Alexandros, Chen, Shizhe, Schmid, Cordelia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
Learning Explicit Contact for Implicit Reconstruction of Hand-held Objects from Monocular Images
by: Hu, Junxing, et al.
Published: (2023)
by: Hu, Junxing, et al.
Published: (2023)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
Online 3D Scene Reconstruction Using Neural Object Priors
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
Design2Cloth: 3D Cloth Generation from 2D Masks
by: Zheng, Jiali, et al.
Published: (2024)
by: Zheng, Jiali, et al.
Published: (2024)
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
by: Si, Chen, et al.
Published: (2026)
by: Si, Chen, et al.
Published: (2026)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation
by: Pratikaki, Chrysa, et al.
Published: (2026)
by: Pratikaki, Chrysa, et al.
Published: (2026)
ShapeFusion: A 3D diffusion model for localized shape editing
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance
by: Gerogiannis, Dimitrios, et al.
Published: (2025)
by: Gerogiannis, Dimitrios, et al.
Published: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024)
by: Ververas, Evangelos, et al.
Published: (2024)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
Locally Adaptive Neural 3D Morphable Models
by: Tarasiou, Michail, et al.
Published: (2024)
by: Tarasiou, Michail, et al.
Published: (2024)
ZeroGS: Training 3D Gaussian Splatting from Unposed Images
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
Neural Sign Actors: A diffusion model for 3D sign language production from text
by: Baltatzis, Vasileios, et al.
Published: (2023)
by: Baltatzis, Vasileios, et al.
Published: (2023)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
AnimateMe: 4D Facial Expressions via Diffusion Models
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions
by: Aytekin, Ayce Idil, et al.
Published: (2026)
by: Aytekin, Ayce Idil, et al.
Published: (2026)
CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-like Contact Representations
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision
by: Zhang, Chenyangguang, et al.
Published: (2023)
by: Zhang, Chenyangguang, et al.
Published: (2023)
HOIST-Former: Hand-held Objects Identification, Segmentation, and Tracking in the Wild
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
by: Lin, Weiquan, et al.
Published: (2026)
by: Lin, Weiquan, et al.
Published: (2026)
TexHOI: Reconstructing Textures of 3D Unknown Objects in Monocular Hand-Object Interaction Scenes
by: Aggarwal, Alakh, et al.
Published: (2025)
by: Aggarwal, Alakh, et al.
Published: (2025)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
by: Fu, Bowen, et al.
Published: (2023)
by: Fu, Bowen, et al.
Published: (2023)
Similar Items
-
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026) -
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025) -
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024) -
Learning Explicit Contact for Implicit Reconstruction of Hand-held Objects from Monocular Images
by: Hu, Junxing, et al.
Published: (2023) -
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)