Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
Fuente:
arXiv
Saved in:
| Main Authors: | Islam, Muhammad, Huang, Tao, Ahn, Euijoon, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review
by: Abdullah, et al.
Published: (2025)
by: Abdullah, et al.
Published: (2025)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
Adversarial Error Correction for Visual Autoregressive Generation
by: Bi, Ligong, et al.
Published: (2026)
by: Bi, Ligong, et al.
Published: (2026)
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions
by: Phuntsho, Karma, et al.
Published: (2025)
by: Phuntsho, Karma, et al.
Published: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
by: Ma, Yiyang, et al.
Published: (2024)
by: Ma, Yiyang, et al.
Published: (2024)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
by: Fu, Xingyu, et al.
Published: (2025)
by: Fu, Xingyu, et al.
Published: (2025)
Generative AI-Driven High-Fidelity Human Motion Simulation
by: Iyer, Hari, et al.
Published: (2025)
by: Iyer, Hari, et al.
Published: (2025)
Contact-aware Human Motion Generation from Textual Descriptions
by: Ma, Sihan, et al.
Published: (2024)
by: Ma, Sihan, et al.
Published: (2024)
KinMo: Kinematic-aware Human Motion Understanding and Generation
by: Zhang, Pengfei, et al.
Published: (2024)
by: Zhang, Pengfei, et al.
Published: (2024)
StickMotion: Generating 3D Human Motions by Drawing a Stickman
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
by: Yang, Xiuyu, et al.
Published: (2025)
by: Yang, Xiuyu, et al.
Published: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
WANDR: Intention-guided Human Motion Generation
by: Diomataris, Markos, et al.
Published: (2024)
by: Diomataris, Markos, et al.
Published: (2024)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
by: Huang, Yidong, et al.
Published: (2026)
by: Huang, Yidong, et al.
Published: (2026)
Generative Pre-trained Autoregressive Diffusion Transformer
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
CapsFake: A Multimodal Capsule Network for Detecting Instruction-Guided Deepfakes
by: Nguyen, Tuan, et al.
Published: (2025)
by: Nguyen, Tuan, et al.
Published: (2025)
MAGI-1: Autoregressive Video Generation at Scale
by: ai, Sand., et al.
Published: (2025)
by: ai, Sand., et al.
Published: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
by: Zhao, Haozhe, et al.
Published: (2025)
by: Zhao, Haozhe, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
by: Chen, Ming, et al.
Published: (2025)
by: Chen, Ming, et al.
Published: (2025)
RU-AI: A Large Multimodal Dataset for Machine-Generated Content Detection
by: Huang, Liting, et al.
Published: (2024)
by: Huang, Liting, et al.
Published: (2024)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
by: Fu, Shuhao, et al.
Published: (2025)
by: Fu, Shuhao, et al.
Published: (2025)
Speculative Decoding for Autoregressive Video Generation
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
Spanning Tree Autoregressive Visual Generation
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
by: Jiang, Yibo, et al.
Published: (2026)
by: Jiang, Yibo, et al.
Published: (2026)
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
by: Zhang, Hong, et al.
Published: (2025)
by: Zhang, Hong, et al.
Published: (2025)
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
by: Cai, Deli, et al.
Published: (2026)
by: Cai, Deli, et al.
Published: (2026)
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
by: Zhang, Zhuoyang, et al.
Published: (2025)
by: Zhang, Zhuoyang, et al.
Published: (2025)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
by: Li, Ruineng, et al.
Published: (2025)
by: Li, Ruineng, et al.
Published: (2025)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Detecting AI-Generated Images via Diffusion Snap-Back Reconstruction: A Forensic Approach
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
Melon Fruit Detection and Quality Assessment Using Generative AI-Based Image Data Augmentation
by: Yoon, Seungri, et al.
Published: (2024)
by: Yoon, Seungri, et al.
Published: (2024)
Similar Items
-
Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review
by: Abdullah, et al.
Published: (2025) -
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023) -
Adversarial Error Correction for Visual Autoregressive Generation
by: Bi, Ligong, et al.
Published: (2026) -
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025) -
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)