TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915647764561920 |
|---|---|
| author | Ma, Yukuo Liu, Cong Wang, Junke Liu, Junqi Huang, Haibin Wu, Zuxuan Zhang, Chi Li, Xuelong |
| author_facet | Ma, Yukuo Liu, Cong Wang, Junke Liu, Junqi Huang, Haibin Wu, Zuxuan Zhang, Chi Li, Xuelong |
| contents | We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_12578 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction Ma, Yukuo Liu, Cong Wang, Junke Liu, Junqi Huang, Haibin Wu, Zuxuan Zhang, Chi Li, Xuelong Computer Vision and Pattern Recognition We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality. |
| title | TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.12578 |