MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zhe, He, Yisheng, Zhong, Lei, Shen, Weichao, Zuo, Qi, Qiu, Lingteng, Dong, Zilong, Yang, Laurence Tianruo, Yuan, Weihao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913741694566400
author Li, Zhe
He, Yisheng
Zhong, Lei
Shen, Weichao
Zuo, Qi
Qiu, Lingteng
Dong, Zilong
Yang, Laurence Tianruo
Yuan, Weihao
author_facet Li, Zhe
He, Yisheng
Zhong, Lei
Shen, Weichao
Zuo, Qi
Qiu, Lingteng
Dong, Zilong
Yang, Laurence Tianruo
Yuan, Weihao
contents Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09901
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
Li, Zhe
He, Yisheng
Zhong, Lei
Shen, Weichao
Zuo, Qi
Qiu, Lingteng
Dong, Zilong
Yang, Laurence Tianruo
Yuan, Weihao
Computer Vision and Pattern Recognition
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
title MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09901