StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yiheng, Yang, Hui, Luo, Chuanchen, Wang, Yuxi, Xu, Shibiao, Zhang, Zhaoxiang, Zhang, Man, Peng, Junran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
by: Zhang, Genghao, et al.
Published: (2024)
by: Zhang, Genghao, et al.
Published: (2024)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
by: Zhang, Xuling, et al.
Published: (2024)
by: Zhang, Xuling, et al.
Published: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
by: Wang, Sen, et al.
Published: (2024)
by: Wang, Sen, et al.
Published: (2024)
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
AMD: Autoregressive Motion Diffusion
by: Han, Bo, et al.
Published: (2023)
by: Han, Bo, et al.
Published: (2023)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction
by: Jiang, Yueheng, et al.
Published: (2025)
by: Jiang, Yueheng, et al.
Published: (2025)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
by: You, Fuming, et al.
Published: (2024)
by: You, Fuming, et al.
Published: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
by: Jin, Chuhao, et al.
Published: (2025)
by: Jin, Chuhao, et al.
Published: (2025)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models
by: Fang, ZhongLi, et al.
Published: (2025)
by: Fang, ZhongLi, et al.
Published: (2025)
Robust Multi-generation Learned Compression of Point Cloud Attribute
by: Liu, Xiangzuo, et al.
Published: (2025)
by: Liu, Xiangzuo, et al.
Published: (2025)
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
StePO-Rec: Towards Personalized Outfit Styling Assistant via Knowledge-Guided Multi-Step Reasoning
by: Bi, Yuxi, et al.
Published: (2025)
by: Bi, Yuxi, et al.
Published: (2025)
PiGW: A Plug-in Generative Watermarking Framework
by: Ma, Rui, et al.
Published: (2024)
by: Ma, Rui, et al.
Published: (2024)
Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise
by: Xu, Zijing, et al.
Published: (2025)
by: Xu, Zijing, et al.
Published: (2025)
Rethinking Security of Diffusion-based Generative Steganography
by: Zhu, Jihao, et al.
Published: (2026)
by: Zhu, Jihao, et al.
Published: (2026)
Rethinking Fusion: Disentangled Learning of Shared and Modality-Specific Information for Stance Detection
by: Xie, Zhiyu, et al.
Published: (2026)
by: Xie, Zhiyu, et al.
Published: (2026)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
by: Zhang, Yiyun, et al.
Published: (2023)
by: Zhang, Yiyun, et al.
Published: (2023)
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
by: Yu, Zihao, et al.
Published: (2025)
by: Yu, Zihao, et al.
Published: (2025)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
by: Li, Qingcao, et al.
Published: (2026)
by: Li, Qingcao, et al.
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
by: Cai, Qi, et al.
Published: (2025)
by: Cai, Qi, et al.
Published: (2025)
SceneX: Procedural Controllable Large-scale Scene Generation
by: Zhou, Mengqi, et al.
Published: (2024)
by: Zhou, Mengqi, et al.
Published: (2024)
CityX: Controllable Procedural Content Generation for Unbounded 3D Cities
by: Zhang, Shougao, et al.
Published: (2024)
by: Zhang, Shougao, et al.
Published: (2024)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
by: Yi, Zijian, et al.
Published: (2024)
by: Yi, Zijian, et al.
Published: (2024)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
Deep Compositional Phase Diffusion for Long Motion Sequence Generation
by: Au, Ho Yin, et al.
Published: (2025)
by: Au, Ho Yin, et al.
Published: (2025)
EV-NVC: Efficient Variable bitrate Neural Video Compression
by: Hu, Yongcun, et al.
Published: (2025)
by: Hu, Yongcun, et al.
Published: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
by: Zhu, Lingsi, et al.
Published: (2026)
by: Zhu, Lingsi, et al.
Published: (2026)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
Similar Items
-
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025) -
FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
by: Zhang, Genghao, et al.
Published: (2024) -
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
by: Zhang, Xuling, et al.
Published: (2024) -
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
by: Wang, Sen, et al.
Published: (2024) -
OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains
by: Zhang, Yixuan, et al.
Published: (2024)