HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Bao, Chen, Xu, Jiarui, Wang, Xiaolong, Gupta, Abhinav, Bharadhwaj, Homanga |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
by: Chen, Hongyi, et al.
Published: (2025)
by: Chen, Hongyi, et al.
Published: (2025)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026)
by: Soraki, Rustin, et al.
Published: (2026)
Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
by: Liu, Mengyuan, et al.
Published: (2024)
by: Liu, Mengyuan, et al.
Published: (2024)
ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
Semantically Controllable Augmentations for Generalizable Robot Learning
by: Chen, Zoey, et al.
Published: (2024)
by: Chen, Zoey, et al.
Published: (2024)
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
by: Chen, Hongyi, et al.
Published: (2026)
by: Chen, Hongyi, et al.
Published: (2026)
TRec: Learning Hand-Object Interactions through 2D Point Track Motion
by: Holzmann, Dennis, et al.
Published: (2026)
by: Holzmann, Dennis, et al.
Published: (2026)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System
by: Qin, Yuzhe, et al.
Published: (2023)
by: Qin, Yuzhe, et al.
Published: (2023)
Learning Continuous Grasping Function with a Dexterous Hand from Human Demonstrations
by: Ye, Jianglong, et al.
Published: (2022)
by: Ye, Jianglong, et al.
Published: (2022)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
by: Yang, Yue, et al.
Published: (2026)
by: Yang, Yue, et al.
Published: (2026)
Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing
by: Yuan, Ying, et al.
Published: (2023)
by: Yuan, Ying, et al.
Published: (2023)
Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
3D Reconstruction of Objects in Hands without Real World 3D Supervision
by: Prakash, Aditya, et al.
Published: (2023)
by: Prakash, Aditya, et al.
Published: (2023)
Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
by: Chen, Mingfei, et al.
Published: (2025)
by: Chen, Mingfei, et al.
Published: (2025)
MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
by: Saleem, Muhammad Usama, et al.
Published: (2024)
by: Saleem, Muhammad Usama, et al.
Published: (2024)
BG-HOP: A Bimanual Generative Hand-Object Prior
by: Krishna, Sriram, et al.
Published: (2025)
by: Krishna, Sriram, et al.
Published: (2025)
DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation
by: Liang, Zhixuan, et al.
Published: (2024)
by: Liang, Zhixuan, et al.
Published: (2024)
HandDGP: Camera-Space Hand Mesh Prediction with Differentiable Global Positioning
by: Valassakis, Eugene, et al.
Published: (2024)
by: Valassakis, Eugene, et al.
Published: (2024)
HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions
by: Xu, Hao, et al.
Published: (2024)
by: Xu, Hao, et al.
Published: (2024)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
by: Liu, Jiajin, et al.
Published: (2026)
by: Liu, Jiajin, et al.
Published: (2026)
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
by: Bimbraw, Keshav, et al.
Published: (2024)
by: Bimbraw, Keshav, et al.
Published: (2024)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions
by: Prakash, Aditya, et al.
Published: (2025)
by: Prakash, Aditya, et al.
Published: (2025)
HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data
by: Zhang, Mengqi, et al.
Published: (2024)
by: Zhang, Mengqi, et al.
Published: (2024)
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
by: Xu, Boshen, et al.
Published: (2024)
by: Xu, Boshen, et al.
Published: (2024)
Walk through Paintings: Egocentric World Models from Internet Priors
by: Bagchi, Anurag, et al.
Published: (2026)
by: Bagchi, Anurag, et al.
Published: (2026)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
by: Chen, Boyuan, et al.
Published: (2024)
by: Chen, Boyuan, et al.
Published: (2024)
SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
by: Xu, Vasco, et al.
Published: (2026)
by: Xu, Vasco, et al.
Published: (2026)
HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models
by: Huang, Mingzhen, et al.
Published: (2025)
by: Huang, Mingzhen, et al.
Published: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
GeneOH Diffusion: Towards Generalizable Hand-Object Interaction Denoising via Denoising Diffusion
by: Liu, Xueyi, et al.
Published: (2024)
by: Liu, Xueyi, et al.
Published: (2024)
Bimanual 3D Hand Motion and Articulation Forecasting in Everyday Images
by: Prakash, Aditya, et al.
Published: (2025)
by: Prakash, Aditya, et al.
Published: (2025)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
by: Gerych, Walter, et al.
Published: (2024)
by: Gerych, Walter, et al.
Published: (2024)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
by: Sivakumar, Anushka, et al.
Published: (2025)
by: Sivakumar, Anushka, et al.
Published: (2025)
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
by: Bansal, Siddhant, et al.
Published: (2024)
by: Bansal, Siddhant, et al.
Published: (2024)
Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting
by: Zhong, Siru, et al.
Published: (2025)
by: Zhong, Siru, et al.
Published: (2025)
Similar Items
-
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
by: Chen, Hongyi, et al.
Published: (2025) -
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024) -
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026) -
Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
by: Liu, Mengyuan, et al.
Published: (2024) -
ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation
by: Zhu, Zehao, et al.
Published: (2023)