StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Ziyu, Lee, Young Yoon, Liu, Joseph, Ben-Shabat, Yizhak, Zordan, Victor, Kapadia, Mubbasir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908288610729984
author Guo, Ziyu
Lee, Young Yoon
Liu, Joseph
Ben-Shabat, Yizhak
Zordan, Victor
Kapadia, Mubbasir
author_facet Guo, Ziyu
Lee, Young Yoon
Liu, Joseph
Ben-Shabat, Yizhak
Zordan, Victor
Kapadia, Mubbasir
contents We present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synthesizes motion across a wide range of content while incorporating stylistic cues from multi-modal inputs, including motion, text, image, video, and audio. To achieve this, we introduce a style-content cross fusion mechanism and align a style encoder with a pre-trained multi-modal model, ensuring that the generated motion accurately captures the reference style while preserving realism. Extensive experiments demonstrate that our framework surpasses existing methods in stylized motion generation and exhibits emergent capabilities for multi-modal motion stylization, enabling more nuanced motion synthesis. Source code and pre-trained models will be released upon acceptance. Project Page: https://stylemotif.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2503_21775
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
Guo, Ziyu
Lee, Young Yoon
Liu, Joseph
Ben-Shabat, Yizhak
Zordan, Victor
Kapadia, Mubbasir
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Graphics
Machine Learning
We present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synthesizes motion across a wide range of content while incorporating stylistic cues from multi-modal inputs, including motion, text, image, video, and audio. To achieve this, we introduce a style-content cross fusion mechanism and align a style encoder with a pre-trained multi-modal model, ensuring that the generated motion accurately captures the reference style while preserving realism. Extensive experiments demonstrate that our framework surpasses existing methods in stylized motion generation and exhibits emergent capabilities for multi-modal motion stylization, enabling more nuanced motion synthesis. Source code and pre-trained models will be released upon acceptance. Project Page: https://stylemotif.github.io
title StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Graphics
Machine Learning
url https://arxiv.org/abs/2503.21775