Text-Driven Video Style Transfer with State-Space Models: Extending StyleMamba for Temporal Coherence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chao, Park, Minsu, Rossi, Cristina, Li, Zhuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916869158469632
author Li, Chao
Park, Minsu
Rossi, Cristina
Li, Zhuang
author_facet Li, Chao
Park, Minsu
Rossi, Cristina
Li, Zhuang
contents StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the StyleMamba framework to handle video sequences. We propose new temporal modules, including a \emph{Video State-Space Fusion Module} to model inter-frame dependencies and a novel \emph{Temporal Masked Directional Loss} that ensures style consistency while addressing scene changes and partial occlusions. Additionally, we introduce a \emph{Temporal Second-Order Loss} to suppress abrupt style variations across consecutive frames. Our experiments on DAVIS and UCF101 show that the proposed approach outperforms competing methods in terms of style consistency, smoothness, and computational efficiency. We believe our new framework paves the way for real-time text-driven video stylization with state-of-the-art perceptual results.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-Driven Video Style Transfer with State-Space Models: Extending StyleMamba for Temporal Coherence
Li, Chao
Park, Minsu
Rossi, Cristina
Li, Zhuang
Graphics
StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the StyleMamba framework to handle video sequences. We propose new temporal modules, including a \emph{Video State-Space Fusion Module} to model inter-frame dependencies and a novel \emph{Temporal Masked Directional Loss} that ensures style consistency while addressing scene changes and partial occlusions. Additionally, we introduce a \emph{Temporal Second-Order Loss} to suppress abrupt style variations across consecutive frames. Our experiments on DAVIS and UCF101 show that the proposed approach outperforms competing methods in terms of style consistency, smoothness, and computational efficiency. We believe our new framework paves the way for real-time text-driven video stylization with state-of-the-art perceptual results.
title Text-Driven Video Style Transfer with State-Space Models: Extending StyleMamba for Temporal Coherence
topic Graphics
url https://arxiv.org/abs/2503.12291