A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Effendy, Edward, Tseng, Kuan-Wei, Kawakami, Rei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model
by: Zhang, Yonghao, et al.
Published: (2024)
by: Zhang, Yonghao, et al.
Published: (2024)
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
by: Xu, Guowei, et al.
Published: (2025)
by: Xu, Guowei, et al.
Published: (2025)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025)
by: An, Zhaoyi, et al.
Published: (2025)
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)
by: Bian, Yuxuan, et al.
Published: (2024)
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
by: Yoshihashi, Ryota, et al.
Published: (2026)
by: Yoshihashi, Ryota, et al.
Published: (2026)
Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts
by: Lee, Eunho, et al.
Published: (2026)
by: Lee, Eunho, et al.
Published: (2026)
UGG: Unified Generative Grasping
by: Lu, Jiaxin, et al.
Published: (2023)
by: Lu, Jiaxin, et al.
Published: (2023)
GraspXL: Generating Grasping Motions for Diverse Objects at Scale
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning
by: Lee, Jihyun, et al.
Published: (2025)
by: Lee, Jihyun, et al.
Published: (2025)
Markerless Motion Capture for Biomechanical Whole-Body Kinematic Estimation in Infants
by: Joshi, Divya, et al.
Published: (2026)
by: Joshi, Divya, et al.
Published: (2026)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion
by: Duran, Enes, et al.
Published: (2026)
by: Duran, Enes, et al.
Published: (2026)
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
by: Tanaka, Daichi, et al.
Published: (2025)
by: Tanaka, Daichi, et al.
Published: (2025)
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation
by: Huang, Zehuan, et al.
Published: (2024)
by: Huang, Zehuan, et al.
Published: (2024)
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
by: Jiang, Haoran, et al.
Published: (2025)
by: Jiang, Haoran, et al.
Published: (2025)
LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
Prototypical Transformer as Unified Motion Learners
by: Han, Cheng, et al.
Published: (2024)
by: Han, Cheng, et al.
Published: (2024)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
by: Ling, Zeyu, et al.
Published: (2024)
by: Ling, Zeyu, et al.
Published: (2024)
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
by: He, Xialin, et al.
Published: (2026)
by: He, Xialin, et al.
Published: (2026)
Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
by: Hsieh, Yi-Kuan, et al.
Published: (2025)
GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction
by: Kwon, Patrick, et al.
Published: (2024)
by: Kwon, Patrick, et al.
Published: (2024)
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
by: Tam, King-Man, et al.
Published: (2025)
by: Tam, King-Man, et al.
Published: (2025)
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
by: Sekikawa, Yusuke, et al.
Published: (2024)
by: Sekikawa, Yusuke, et al.
Published: (2024)
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
by: Kada, Masahiro, et al.
Published: (2026)
by: Kada, Masahiro, et al.
Published: (2026)
RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
by: Chen, Jiaben, et al.
Published: (2024)
by: Chen, Jiaben, et al.
Published: (2024)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
by: Niu, Yida, et al.
Published: (2026)
by: Niu, Yida, et al.
Published: (2026)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
by: Chu, Xiangxiang, et al.
Published: (2025)
by: Chu, Xiangxiang, et al.
Published: (2025)
GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation
by: Shi, Junyu, et al.
Published: (2025)
by: Shi, Junyu, et al.
Published: (2025)
3D Whole-body Grasp Synthesis with Directional Controllability
by: Paschalidis, Georgios, et al.
Published: (2024)
by: Paschalidis, Georgios, et al.
Published: (2024)
Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation
by: Zhu, Vince, et al.
Published: (2024)
by: Zhu, Vince, et al.
Published: (2024)
Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
by: Wang, Chaoyi, et al.
Published: (2025)
by: Wang, Chaoyi, et al.
Published: (2025)
Large Motion Model for Unified Multi-Modal Motion Generation
by: Zhang, Mingyuan, et al.
Published: (2024)
by: Zhang, Mingyuan, et al.
Published: (2024)
SpeechAct: Towards Generating Whole-body Motion from Speech
by: Zhang, Jinsong, et al.
Published: (2023)
by: Zhang, Jinsong, et al.
Published: (2023)
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
by: Chen, Bohong, et al.
Published: (2024)
by: Chen, Bohong, et al.
Published: (2024)
MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
by: Guo, Ziyan, et al.
Published: (2025)
by: Guo, Ziyan, et al.
Published: (2025)
Similar Items
-
Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model
by: Zhang, Yonghao, et al.
Published: (2024) -
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
by: Xu, Guowei, et al.
Published: (2025) -
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
by: An, Zhaoyi, et al.
Published: (2025) -
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024) -
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
by: Yoshihashi, Ryota, et al.
Published: (2026)