M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chi, Seunggeun, Chi, Hyung-gun, Ma, Hengbo, Agarwal, Nakul, Siddiqui, Faizan, Ramani, Karthik, Lee, Kwonjoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfoGCN++: Learning Representation by Predicting the Future for Online Human Skeleton-based Action Recognition
by: Chi, Seunggeun, et al.
Published: (2023)
by: Chi, Seunggeun, et al.
Published: (2023)
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data
by: Chi, Seunggeun, et al.
Published: (2024)
by: Chi, Seunggeun, et al.
Published: (2024)
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
by: Doh, Hyungjun, et al.
Published: (2025)
by: Doh, Hyungjun, et al.
Published: (2025)
Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
by: Lee, Dong In, et al.
Published: (2025)
by: Lee, Dong In, et al.
Published: (2025)
Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional Inpainting
by: Chi, Seunggeun, et al.
Published: (2025)
by: Chi, Seunggeun, et al.
Published: (2025)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
by: Mittal, Himangi, et al.
Published: (2024)
by: Mittal, Himangi, et al.
Published: (2024)
CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
by: Roh, Wonseok, et al.
Published: (2024)
by: Roh, Wonseok, et al.
Published: (2024)
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023)
by: Wang, Shijie, et al.
Published: (2023)
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
by: Jang, Youngdong, et al.
Published: (2026)
by: Jang, Youngdong, et al.
Published: (2026)
AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?
by: Zhao, Qi, et al.
Published: (2023)
by: Zhao, Qi, et al.
Published: (2023)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
by: Chi, Xiaowei, et al.
Published: (2023)
by: Chi, Xiaowei, et al.
Published: (2023)
DYNAMO: Dependency-Aware Deep Learning Framework for Articulated Assembly Motion Prediction
by: Patel, Mayank, et al.
Published: (2025)
by: Patel, Mayank, et al.
Published: (2025)
AToM: Amortized Text-to-Mesh using 2D Diffusion
by: Qian, Guocheng, et al.
Published: (2024)
by: Qian, Guocheng, et al.
Published: (2024)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
by: Oh, Gyeongrok, et al.
Published: (2025)
by: Oh, Gyeongrok, et al.
Published: (2025)
T2M Mamba: Motion Periodicity-Saliency Coupling Approach for Stable Text-Driven Motion Generation
by: Zhan, Xingzu, et al.
Published: (2026)
by: Zhan, Xingzu, et al.
Published: (2026)
GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation
by: Shi, Junyu, et al.
Published: (2025)
by: Shi, Junyu, et al.
Published: (2025)
Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
by: Moon, Seokha, et al.
Published: (2024)
by: Moon, Seokha, et al.
Published: (2024)
M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2026)
by: Symeonidis-Herzig, Alexandre, et al.
Published: (2026)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain
by: Chen, Wenshuo, et al.
Published: (2025)
by: Chen, Wenshuo, et al.
Published: (2025)
T2M-X: Learning Expressive Text-to-Motion Generation from Partially Annotated Data
by: Liu, Mingdian, et al.
Published: (2024)
by: Liu, Mingdian, et al.
Published: (2024)
Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
by: Wang, Yin, et al.
Published: (2025)
by: Wang, Yin, et al.
Published: (2025)
PersonaBooth: Personalized Text-to-Motion Generation
by: Kim, Boeun, et al.
Published: (2025)
by: Kim, Boeun, et al.
Published: (2025)
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
by: Sharma, Nakul, et al.
Published: (2024)
by: Sharma, Nakul, et al.
Published: (2024)
DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
by: Zhang, Ning, et al.
Published: (2026)
by: Zhang, Ning, et al.
Published: (2026)
Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-view Images
by: Toso, Matteo, et al.
Published: (2024)
by: Toso, Matteo, et al.
Published: (2024)
MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion
by: Susladkar, Onkar, et al.
Published: (2024)
by: Susladkar, Onkar, et al.
Published: (2024)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
by: Chen, Xingyu
Published: (2024)
by: Chen, Xingyu
Published: (2024)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
by: Agarwal, Nakul, et al.
Published: (2026)
by: Agarwal, Nakul, et al.
Published: (2026)
Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion
by: Lu, Yuanxun, et al.
Published: (2023)
by: Lu, Yuanxun, et al.
Published: (2023)
Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction
by: Cha, Junuk, et al.
Published: (2024)
by: Cha, Junuk, et al.
Published: (2024)
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
by: Liu, Yifei, et al.
Published: (2026)
by: Liu, Yifei, et al.
Published: (2026)
Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation
by: Cheng, Ching-Lam, et al.
Published: (2026)
by: Cheng, Ching-Lam, et al.
Published: (2026)
CMP: Cooperative Motion Prediction with Multi-Agent Communication
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
T3M: Text Guided 3D Human Motion Synthesis from Speech
by: Peng, Wenshuo, et al.
Published: (2024)
by: Peng, Wenshuo, et al.
Published: (2024)
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
by: Ghoddoosian, Reza, et al.
Published: (2024)
by: Ghoddoosian, Reza, et al.
Published: (2024)
Similar Items
-
InfoGCN++: Learning Representation by Predicting the Future for Online Human Skeleton-based Action Recognition
by: Chi, Seunggeun, et al.
Published: (2023) -
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data
by: Chi, Seunggeun, et al.
Published: (2024) -
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
by: Doh, Hyungjun, et al.
Published: (2025) -
Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformer
by: Lee, Dong In, et al.
Published: (2025) -
Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional Inpainting
by: Chi, Seunggeun, et al.
Published: (2025)