TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Yukuo, Liu, Cong, Wang, Junke, Liu, Junqi, Huang, Haibin, Wu, Zuxuan, Zhang, Chi, Li, Xuelong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915647764561920
author Ma, Yukuo
Liu, Cong
Wang, Junke
Liu, Junqi
Huang, Haibin
Wu, Zuxuan
Zhang, Chi
Li, Xuelong
author_facet Ma, Yukuo
Liu, Cong
Wang, Junke
Liu, Junqi
Huang, Haibin
Wu, Zuxuan
Zhang, Chi
Li, Xuelong
contents We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
Ma, Yukuo
Liu, Cong
Wang, Junke
Liu, Junqi
Huang, Haibin
Wu, Zuxuan
Zhang, Chi
Li, Xuelong
Computer Vision and Pattern Recognition
We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality.
title TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12578