VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yiren, Yao, Wangzi, Wang, Haofan, Shou, Mike Zheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916020562690048
author Song, Yiren
Yao, Wangzi
Wang, Haofan
Shou, Mike Zheng
author_facet Song, Yiren
Yao, Wangzi
Wang, Haofan
Shou, Mike Zheng
contents Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylization remains challenging due to temporal inconsistency. Most existing methods stylize frames or keyframes and enforce consistency via heuristic temporal propagation, which is brittle under occlusions, disocclusions, and long-term motion, leading to drift and flickering artifacts. We argue that a fundamental bottleneck lies in the lack of large-scale triplet data and a principled training paradigm that jointly models and disentangles style, content, and motion.To address this, we introduce VISTA-1000, a synthetic dataset with 1,000 styles and motion-aligned triplets of style reference, clean video, and stylized video, and propose a diffusion-transformer-based in-context video style transfer framework with a lightweight style adapter for robust style extraction. Extensive experiments demonstrate SOTA performance in style fidelity, temporal consistency, and content preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
Song, Yiren
Yao, Wangzi
Wang, Haofan
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylization remains challenging due to temporal inconsistency. Most existing methods stylize frames or keyframes and enforce consistency via heuristic temporal propagation, which is brittle under occlusions, disocclusions, and long-term motion, leading to drift and flickering artifacts. We argue that a fundamental bottleneck lies in the lack of large-scale triplet data and a principled training paradigm that jointly models and disentangles style, content, and motion.To address this, we introduce VISTA-1000, a synthetic dataset with 1,000 styles and motion-aligned triplets of style reference, clean video, and stylized video, and propose a diffusion-transformer-based in-context video style transfer framework with a lightweight style adapter for robust style extraction. Extensive experiments demonstrate SOTA performance in style fidelity, temporal consistency, and content preservation.
title VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17312