TeleStyle: Content-Preserving Style Transfer in Images and Videos

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Shiwen, Yang, Xiaoyan, Zi, Bojia, Huang, Haibin, Zhang, Chi, Li, Xuelong
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910003164610560
author Zhang, Shiwen
Yang, Xiaoyan
Zi, Bojia
Huang, Haibin
Zhang, Chi
Li, Xuelong
author_facet Zhang, Shiwen
Yang, Xiaoyan
Zi, Bojia
Huang, Haibin
Zhang, Chi
Li, Xuelong
contents Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their internal representations. In this technical report, we present TeleStyle, a lightweight yet effective model for both image and video stylization. Built upon Qwen-Image-Edit, TeleStyle leverages the base model's robust capabilities in content preservation and style customization. To facilitate effective training, we curated a high-quality dataset of distinct specific styles and further synthesized triplets using thousands of diverse, in-the-wild style categories. We introduce a Curriculum Continual Learning framework to train TeleStyle on this hybrid dataset of clean (curated) and noisy (synthetic) triplets. This approach enables the model to generalize to unseen styles without compromising precise content fidelity. Additionally, we introduce a video-to-video stylization module to enhance temporal consistency and visual quality. TeleStyle achieves state-of-the-art performance across three core evaluation metrics: style similarity, content consistency, and aesthetic quality. Code and pre-trained models are available at https://github.com/Tele-AI/TeleStyle
format Preprint
id arxiv_https___arxiv_org_abs_2601_20175
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TeleStyle: Content-Preserving Style Transfer in Images and Videos
Zhang, Shiwen
Yang, Xiaoyan
Zi, Bojia
Huang, Haibin
Zhang, Chi
Li, Xuelong
Computer Vision and Pattern Recognition
Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their internal representations. In this technical report, we present TeleStyle, a lightweight yet effective model for both image and video stylization. Built upon Qwen-Image-Edit, TeleStyle leverages the base model's robust capabilities in content preservation and style customization. To facilitate effective training, we curated a high-quality dataset of distinct specific styles and further synthesized triplets using thousands of diverse, in-the-wild style categories. We introduce a Curriculum Continual Learning framework to train TeleStyle on this hybrid dataset of clean (curated) and noisy (synthetic) triplets. This approach enables the model to generalize to unseen styles without compromising precise content fidelity. Additionally, we introduce a video-to-video stylization module to enhance temporal consistency and visual quality. TeleStyle achieves state-of-the-art performance across three core evaluation metrics: style similarity, content consistency, and aesthetic quality. Code and pre-trained models are available at https://github.com/Tele-AI/TeleStyle
title TeleStyle: Content-Preserving Style Transfer in Images and Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.20175