Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917147318419456 |
|---|---|
| author | Zhang, Jiangning Zhu, Junwei Hu, Teng Wang, Yabiao Luo, Donghao Cao, Weijian Gan, Zhenye Hu, Xiaobin Xue, Zhucun Wang, Chengjie |
| author_facet | Zhang, Jiangning Zhu, Junwei Hu, Teng Wang, Yabiao Luo, Donghao Cao, Weijian Gan, Zhenye Hu, Xiaobin Xue, Zhucun Wang, Chengjie |
| contents | Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed $\textbf{T3}$ ($\textbf{T}$ransform $\textbf{T}$rained $\textbf{T}$ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, $\textbf{T3-Video}$ introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that $\textbf{T3-Video}$ substantially outperforms existing approaches: while delivering performance improvements (+4.29$\uparrow$ VQA and +0.08$\uparrow$ VTC), it accelerates native 4K video generation by more than 10$\times$. Project page at https://zhangzjn.github.io/projects/T3-Video |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_13492 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$ Zhang, Jiangning Zhu, Junwei Hu, Teng Wang, Yabiao Luo, Donghao Cao, Weijian Gan, Zhenye Hu, Xiaobin Xue, Zhucun Wang, Chengjie Computer Vision and Pattern Recognition Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed $\textbf{T3}$ ($\textbf{T}$ransform $\textbf{T}$rained $\textbf{T}$ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, $\textbf{T3-Video}$ introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that $\textbf{T3-Video}$ substantially outperforms existing approaches: while delivering performance improvements (+4.29$\uparrow$ VQA and +0.08$\uparrow$ VTC), it accelerates native 4K video generation by more than 10$\times$. Project page at https://zhangzjn.github.io/projects/T3-Video |
| title | Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$ |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.13492 |