VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yiren, Yao, Wangzi, Wang, Haofan, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniPSD: Layered PSD Generation with Diffusion Transformer
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
by: Song, Yiren, et al.
Published: (2024)
by: Song, Yiren, et al.
Published: (2024)
Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
by: Shi, Yiqing, et al.
Published: (2025)
by: Shi, Yiqing, et al.
Published: (2025)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025)
by: Lu, Runnan, et al.
Published: (2025)
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
by: Zhang, Yuxuan, et al.
Published: (2025)
by: Zhang, Yuxuan, et al.
Published: (2025)
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
by: Zheng, Yiren, et al.
Published: (2026)
by: Zheng, Yiren, et al.
Published: (2026)
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
by: Yang, Pei, et al.
Published: (2026)
by: Yang, Pei, et al.
Published: (2026)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious?
by: Yang, Pei, et al.
Published: (2024)
by: Yang, Pei, et al.
Published: (2024)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
by: Song, Yiren, et al.
Published: (2024)
by: Song, Yiren, et al.
Published: (2024)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
by: Ci, Hai, et al.
Published: (2024)
by: Ci, Hai, et al.
Published: (2024)
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
by: Ci, Hai, et al.
Published: (2025)
by: Ci, Hai, et al.
Published: (2025)
Image Watermarks are Removable Using Controllable Regeneration from Clean Noise
by: Liu, Yepeng, et al.
Published: (2024)
by: Liu, Yepeng, et al.
Published: (2024)
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
by: Ci, Hai, et al.
Published: (2024)
by: Ci, Hai, et al.
Published: (2024)
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
by: Zhang, Xiaoyan, et al.
Published: (2026)
by: Zhang, Xiaoyan, et al.
Published: (2026)
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
by: Gao, Heyuan, et al.
Published: (2026)
by: Gao, Heyuan, et al.
Published: (2026)
InstantStyle-Plus: Style Transfer with Content-Preserving in Text-to-Image Generation
by: Wang, Haofan, et al.
Published: (2024)
by: Wang, Haofan, et al.
Published: (2024)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
D-AR: Diffusion via Autoregressive Models
by: Gao, Ziteng, et al.
Published: (2025)
by: Gao, Ziteng, et al.
Published: (2025)
Loom: Diffusion-Transformer for Interleaved Generation
by: Ye, Mingcheng, et al.
Published: (2025)
by: Ye, Mingcheng, et al.
Published: (2025)
Impossible Videos
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
TransAnimate: Taming Layer Diffusion to Generate RGBA Video
by: Chen, Xuewei, et al.
Published: (2025)
by: Chen, Xuewei, et al.
Published: (2025)
Diffusion-Driven Self-Supervised Learning for Shape Reconstruction and Pose Estimation
by: Sun, Jingtao, et al.
Published: (2024)
by: Sun, Jingtao, et al.
Published: (2024)
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Single Trajectory Distillation for Accelerating Image and Video Style Transfer
by: Xu, Sijie, et al.
Published: (2024)
by: Xu, Sijie, et al.
Published: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data
by: Huang, Shijie, et al.
Published: (2025)
by: Huang, Shijie, et al.
Published: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
by: Mehraban, Soroush, et al.
Published: (2025)
by: Mehraban, Soroush, et al.
Published: (2025)
Similar Items
-
OmniPSD: Layered PSD Generation with Diffusion Transformer
by: Liu, Cheng, et al.
Published: (2025) -
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
by: Song, Yiren, et al.
Published: (2025) -
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
by: Song, Yiren, et al.
Published: (2025) -
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
by: Song, Yiren, et al.
Published: (2026) -
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
by: Song, Yiren, et al.
Published: (2025)