MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Yiming, Ji, Wei, Zheng, Kecheng, Wang, Zicheng, Xu, Dong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
por: Xu, Haidong, et al.
Publicado: (2025)
por: Xu, Haidong, et al.
Publicado: (2025)
BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation
por: Hosseyni, S. Rohollah, et al.
Publicado: (2024)
por: Hosseyni, S. Rohollah, et al.
Publicado: (2024)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
por: Yang, Mingyue, et al.
Publicado: (2025)
por: Yang, Mingyue, et al.
Publicado: (2025)
Optimizing Prompts for Text-to-Image Generation
por: Hao, Yaru, et al.
Publicado: (2022)
por: Hao, Yaru, et al.
Publicado: (2022)
Scaling Concept With Text-Guided Diffusion Models
por: Huang, Chao, et al.
Publicado: (2024)
por: Huang, Chao, et al.
Publicado: (2024)
Bilingual Text-to-Motion Generation: A New Benchmark and Baselines
por: Weng, Wanjiang, et al.
Publicado: (2026)
por: Weng, Wanjiang, et al.
Publicado: (2026)
DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
por: Zhang, Ning, et al.
Publicado: (2026)
por: Zhang, Ning, et al.
Publicado: (2026)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
por: Jia, Mengzhao, et al.
Publicado: (2024)
por: Jia, Mengzhao, et al.
Publicado: (2024)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
por: Li, Xiaomin, et al.
Publicado: (2024)
por: Li, Xiaomin, et al.
Publicado: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
SegMo: Segment-aligned Text to 3D Human Motion Generation
por: Dang, Bowen, et al.
Publicado: (2025)
por: Dang, Bowen, et al.
Publicado: (2025)
Learning Visual Generative Priors without Text
por: Ma, Shuailei, et al.
Publicado: (2024)
por: Ma, Shuailei, et al.
Publicado: (2024)
Mojito: Motion Trajectory and Intensity Control for Video Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
por: Shi, Shuwei, et al.
Publicado: (2024)
por: Shi, Shuwei, et al.
Publicado: (2024)
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
por: Wan, Yixin, et al.
Publicado: (2025)
por: Wan, Yixin, et al.
Publicado: (2025)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
por: Wang, Chenglin, et al.
Publicado: (2026)
por: Wang, Chenglin, et al.
Publicado: (2026)
Image-to-LaTeX Converter for Mathematical Formulas and Text
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
MoExtend: Tuning New Experts for Modality and Task Extension
por: Zhong, Shanshan, et al.
Publicado: (2024)
por: Zhong, Shanshan, et al.
Publicado: (2024)
MoLingo: Motion-Language Alignment for Text-to-Motion Generation
por: He, Yannan, et al.
Publicado: (2025)
por: He, Yannan, et al.
Publicado: (2025)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
por: Chin, Zhi-Yi, et al.
Publicado: (2023)
Progressive Human Motion Generation Based on Text and Few Motion Frames
por: Zeng, Ling-An, et al.
Publicado: (2025)
por: Zeng, Ling-An, et al.
Publicado: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
por: Yuan, Huizhuo, et al.
Publicado: (2024)
por: Yuan, Huizhuo, et al.
Publicado: (2024)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
por: Du, Shian, et al.
Publicado: (2024)
por: Du, Shian, et al.
Publicado: (2024)
Improving OCR for Historical Texts of Multiple Languages
por: Westerdijk, Hylke, et al.
Publicado: (2025)
por: Westerdijk, Hylke, et al.
Publicado: (2025)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
por: Zhao, Qinyu, et al.
Publicado: (2025)
por: Zhao, Qinyu, et al.
Publicado: (2025)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
por: Dong, Ziyi, et al.
Publicado: (2022)
por: Dong, Ziyi, et al.
Publicado: (2022)
You Think, You ACT: The New Task of Arbitrary Text to Motion Generation
por: Wang, Runqi, et al.
Publicado: (2024)
por: Wang, Runqi, et al.
Publicado: (2024)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
por: Jiang, Yue, et al.
Publicado: (2026)
por: Jiang, Yue, et al.
Publicado: (2026)
LaMoD: Latent Motion Diffusion Model For Myocardial Strain Generation
por: Xing, Jiarui, et al.
Publicado: (2024)
por: Xing, Jiarui, et al.
Publicado: (2024)
CoMo: Compositional Motion Customization for Text-to-Video Generation
por: Xu, Youcan, et al.
Publicado: (2025)
por: Xu, Youcan, et al.
Publicado: (2025)
Multimedia Generative Script Learning for Task Planning
por: Wang, Qingyun, et al.
Publicado: (2022)
por: Wang, Qingyun, et al.
Publicado: (2022)
Can We Predict Performance of Large Models across Vision-Language Tasks?
por: Zhao, Qinyu, et al.
Publicado: (2024)
por: Zhao, Qinyu, et al.
Publicado: (2024)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
por: Chen, Zeren, et al.
Publicado: (2023)
por: Chen, Zeren, et al.
Publicado: (2023)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
por: Tan, Shuai, et al.
Publicado: (2024)
por: Tan, Shuai, et al.
Publicado: (2024)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
por: Chen, Xingyu
Publicado: (2024)
por: Chen, Xingyu
Publicado: (2024)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
por: Wang, Yuchi, et al.
Publicado: (2024)
por: Wang, Yuchi, et al.
Publicado: (2024)
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
por: Yan, Sheng, et al.
Publicado: (2023)
por: Yan, Sheng, et al.
Publicado: (2023)
Why Instruction-Based Unlearning Fails in Diffusion Models?
por: Zhang, Zeliang, et al.
Publicado: (2026)
por: Zhang, Zeliang, et al.
Publicado: (2026)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
por: Dang, Lingwei, et al.
Publicado: (2025)
por: Dang, Lingwei, et al.
Publicado: (2025)
Ejemplares similares
-
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
por: Xu, Haidong, et al.
Publicado: (2025) -
BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation
por: Hosseyni, S. Rohollah, et al.
Publicado: (2024) -
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
por: Yang, Mingyue, et al.
Publicado: (2025) -
Optimizing Prompts for Text-to-Image Generation
por: Hao, Yaru, et al.
Publicado: (2022) -
Scaling Concept With Text-Guided Diffusion Models
por: Huang, Chao, et al.
Publicado: (2024)