HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Mingzhen, Chu, Fu-Jen, Tekin, Bugra, Liang, Kevin J, Ma, Haoyu, Wang, Weiyao, Chen, Xingyu, Gleize, Pierre, Xue, Hongfei, Lyu, Siwei, Kitani, Kris, Feiszli, Matt, Tang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
ICON: Incremental CONfidence for Joint Pose and Radiance Field Optimization
by: Wang, Weiyao, et al.
Published: (2024)
by: Wang, Weiyao, et al.
Published: (2024)
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB
by: Lin, Yunzhi, et al.
Published: (2024)
by: Lin, Yunzhi, et al.
Published: (2024)
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024)
by: Ye, Yufei, et al.
Published: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)
by: Xu, Runsen, et al.
Published: (2025)
Joint Diffusion for Universal Hand-Object Grasp Generation
by: Cao, Jinkun, et al.
Published: (2024)
by: Cao, Jinkun, et al.
Published: (2024)
3x2: 3D Object Part Segmentation by 2D Semantic Correspondences
by: Thai, Anh, et al.
Published: (2024)
by: Thai, Anh, et al.
Published: (2024)
Multi-Object Tracking by Hierarchical Visual Representations
by: Cao, Jinkun, et al.
Published: (2024)
by: Cao, Jinkun, et al.
Published: (2024)
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024)
by: Christen, Sammy, et al.
Published: (2024)
Evaluating a VR System for Collecting Safety-Critical Vehicle-Pedestrian Interactions
by: Weng, Erica, et al.
Published: (2023)
by: Weng, Erica, et al.
Published: (2023)
JaywalkerVR: A VR System for Collecting Safety-Critical Pedestrian-Vehicle Interactions
by: Mukoya, Kenta, et al.
Published: (2024)
by: Mukoya, Kenta, et al.
Published: (2024)
Your Text Encoder Can Be An Object-Level Watermarking Controller
by: Devulapally, Naresh Kumar, et al.
Published: (2025)
by: Devulapally, Naresh Kumar, et al.
Published: (2025)
SAM 3D: 3Dfy Anything in Images
by: SAM 3D Team, et al.
Published: (2025)
by: SAM 3D Team, et al.
Published: (2025)
Omnigrasp: Grasping Diverse Objects with Simulated Humanoids
by: Luo, Zhengyi, et al.
Published: (2024)
by: Luo, Zhengyi, et al.
Published: (2024)
Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025)
by: Fan, Zicong, et al.
Published: (2025)
Harmony4D: A Video Dataset for In-The-Wild Close Human Interactions
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
Les chiens s´approchent, et s´éloignent
by: Jean-Marie Gleize
Published: (2007)
by: Jean-Marie Gleize
Published: (2007)
Hand-Object Interaction Controller (HOIC): Deep Reinforcement Learning for Reconstructing Interactions with Physics
by: Hu, Haoyu, et al.
Published: (2024)
by: Hu, Haoyu, et al.
Published: (2024)
Zero-Shot Multi-Object Scene Completion
by: Iwase, Shun, et al.
Published: (2024)
by: Iwase, Shun, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
by: Dou, Zi-Yi, et al.
Published: (2024)
by: Dou, Zi-Yi, et al.
Published: (2024)
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
by: Fang, Qihang, et al.
Published: (2024)
by: Fang, Qihang, et al.
Published: (2024)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
by: Li, Zisu, et al.
Published: (2025)
by: Li, Zisu, et al.
Published: (2025)
CacheFlow: Fast Human Motion Prediction by Cached Normalizing Flow
by: Maeda, Takahiro, et al.
Published: (2025)
by: Maeda, Takahiro, et al.
Published: (2025)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
by: Newman, Benjamin A., et al.
Published: (2024)
by: Newman, Benjamin A., et al.
Published: (2024)
SAM 3D Body: Robust Full-Body Human Mesh Recovery
by: Yang, Xitong, et al.
Published: (2026)
by: Yang, Xitong, et al.
Published: (2026)
ParallelEdits: Efficient Multi-object Image Editing
by: Huang, Mingzhen, et al.
Published: (2024)
by: Huang, Mingzhen, et al.
Published: (2024)
BiDexHand: Design and Evaluation of an Open-Source 16-DoF Biomimetic Dexterous Hand
by: Weng, Zhengyang Kris
Published: (2025)
by: Weng, Zhengyang Kris
Published: (2025)
FoundPose: Unseen Object Pose Estimation with Foundation Features
by: Örnek, Evin Pınar, et al.
Published: (2023)
by: Örnek, Evin Pınar, et al.
Published: (2023)
Bootstrapping Linear Models for Fast Online Adaptation in Human-Agent Collaboration
by: Newman, Benjamin A, et al.
Published: (2024)
by: Newman, Benjamin A, et al.
Published: (2024)
GoTrack: Generic 6DoF Object Pose Refinement and Tracking
by: Nguyen, Van Nguyen, et al.
Published: (2025)
by: Nguyen, Van Nguyen, et al.
Published: (2025)
Exposing Text-Image Inconsistency Using Diffusion Models
by: Huang, Mingzhen, et al.
Published: (2024)
by: Huang, Mingzhen, et al.
Published: (2024)
Hierarchical Procedural Framework for Low-latency Robot-Assisted Hand-Object Interaction
by: Yuan, Mingqi, et al.
Published: (2024)
by: Yuan, Mingqi, et al.
Published: (2024)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image
by: Ma, Xiaoxuan, et al.
Published: (2026)
by: Ma, Xiaoxuan, et al.
Published: (2026)
MGF: Mixed Gaussian Flow for Diverse Trajectory Prediction
by: Chen, Jiahe, et al.
Published: (2024)
by: Chen, Jiahe, et al.
Published: (2024)
ExpertAF: Expert Actionable Feedback from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Generalizable Neural Human Renderer
by: Masuda, Mana, et al.
Published: (2024)
by: Masuda, Mana, et al.
Published: (2024)
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
by: Movahed, Yousef Azizi, et al.
Published: (2025)
by: Movahed, Yousef Azizi, et al.
Published: (2025)
Similar Items
-
ADen: Adaptive Density Representations for Sparse-view Camera Pose Estimation
by: Tang, Hao, et al.
Published: (2024) -
ICON: Incremental CONfidence for Joint Pose and Radiance Field Optimization
by: Wang, Weiyao, et al.
Published: (2024) -
OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB
by: Lin, Yunzhi, et al.
Published: (2024) -
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
by: Ye, Yufei, et al.
Published: (2024) -
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
by: Xu, Runsen, et al.
Published: (2025)