HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Mingzhen, Chu, Fu-Jen, Tekin, Bugra, Liang, Kevin J, Ma, Haoyu, Wang, Weiyao, Chen, Xingyu, Gleize, Pierre, Xue, Hongfei, Lyu, Siwei, Kitani, Kris, Feiszli, Matt, Tang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917967797682176
author Huang, Mingzhen
Chu, Fu-Jen
Tekin, Bugra
Liang, Kevin J
Ma, Haoyu
Wang, Weiyao
Chen, Xingyu
Gleize, Pierre
Xue, Hongfei
Lyu, Siwei
Kitani, Kris
Feiszli, Matt
Tang, Hao
author_facet Huang, Mingzhen
Chu, Fu-Jen
Tekin, Bugra
Liang, Kevin J
Ma, Haoyu
Wang, Weiyao
Chen, Xingyu
Gleize, Pierre
Xue, Hongfei
Lyu, Siwei
Kitani, Kris
Feiszli, Matt
Tang, Hao
contents We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (\eg text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models
Huang, Mingzhen
Chu, Fu-Jen
Tekin, Bugra
Liang, Kevin J
Ma, Haoyu
Wang, Weiyao
Chen, Xingyu
Gleize, Pierre
Xue, Hongfei
Lyu, Siwei
Kitani, Kris
Feiszli, Matt
Tang, Hao
Computer Vision and Pattern Recognition
We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (\eg text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks.
title HOIGPT: Learning Long Sequence Hand-Object Interaction with Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19157