A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913921260060672 |
|---|---|
| author | Effendy, Edward Tseng, Kuan-Wei Kawakami, Rei |
| author_facet | Effendy, Edward Tseng, Kuan-Wei Kawakami, Rei |
| contents | Accepted in the ICIP 2025
We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_00676 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation Effendy, Edward Tseng, Kuan-Wei Kawakami, Rei Computer Vision and Pattern Recognition Accepted in the ICIP 2025 We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications. |
| title | A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2507.00676 |