A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Effendy, Edward, Tseng, Kuan-Wei, Kawakami, Rei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913921260060672
author Effendy, Edward
Tseng, Kuan-Wei
Kawakami, Rei
author_facet Effendy, Edward
Tseng, Kuan-Wei
Kawakami, Rei
contents Accepted in the ICIP 2025 We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
Effendy, Edward
Tseng, Kuan-Wei
Kawakami, Rei
Computer Vision and Pattern Recognition
Accepted in the ICIP 2025 We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications.
title A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.00676