HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917967797682176 |
|---|---|
| author | Huang, Mingzhen Chu, Fu-Jen Tekin, Bugra Liang, Kevin J Ma, Haoyu Wang, Weiyao Chen, Xingyu Gleize, Pierre Xue, Hongfei Lyu, Siwei Kitani, Kris Feiszli, Matt Tang, Hao |
| author_facet | Huang, Mingzhen Chu, Fu-Jen Tekin, Bugra Liang, Kevin J Ma, Haoyu Wang, Weiyao Chen, Xingyu Gleize, Pierre Xue, Hongfei Lyu, Siwei Kitani, Kris Feiszli, Matt Tang, Hao |
| contents | We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (\eg text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_19157 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models Huang, Mingzhen Chu, Fu-Jen Tekin, Bugra Liang, Kevin J Ma, Haoyu Wang, Weiyao Chen, Xingyu Gleize, Pierre Xue, Hongfei Lyu, Siwei Kitani, Kris Feiszli, Matt Tang, Hao Computer Vision and Pattern Recognition We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (\eg text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks. |
| title | HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.19157 |