Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jiangning, Zhu, Junwei, Hu, Teng, Wang, Yabiao, Luo, Donghao, Cao, Weijian, Gan, Zhenye, Hu, Xiaobin, Xue, Zhucun, Wang, Chengjie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917147318419456
author Zhang, Jiangning
Zhu, Junwei
Hu, Teng
Wang, Yabiao
Luo, Donghao
Cao, Weijian
Gan, Zhenye
Hu, Xiaobin
Xue, Zhucun
Wang, Chengjie
author_facet Zhang, Jiangning
Zhu, Junwei
Hu, Teng
Wang, Yabiao
Luo, Donghao
Cao, Weijian
Gan, Zhenye
Hu, Xiaobin
Xue, Zhucun
Wang, Chengjie
contents Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed $\textbf{T3}$ ($\textbf{T}$ransform $\textbf{T}$rained $\textbf{T}$ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, $\textbf{T3-Video}$ introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that $\textbf{T3-Video}$ substantially outperforms existing approaches: while delivering performance improvements (+4.29$\uparrow$ VQA and +0.08$\uparrow$ VTC), it accelerates native 4K video generation by more than 10$\times$. Project page at https://zhangzjn.github.io/projects/T3-Video
format Preprint
id arxiv_https___arxiv_org_abs_2512_13492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
Zhang, Jiangning
Zhu, Junwei
Hu, Teng
Wang, Yabiao
Luo, Donghao
Cao, Weijian
Gan, Zhenye
Hu, Xiaobin
Xue, Zhucun
Wang, Chengjie
Computer Vision and Pattern Recognition
Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed $\textbf{T3}$ ($\textbf{T}$ransform $\textbf{T}$rained $\textbf{T}$ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, $\textbf{T3-Video}$ introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that $\textbf{T3-Video}$ substantially outperforms existing approaches: while delivering performance improvements (+4.29$\uparrow$ VQA and +0.08$\uparrow$ VTC), it accelerates native 4K video generation by more than 10$\times$. Project page at https://zhangzjn.github.io/projects/T3-Video
title Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13492