Video-GPT via Next Clip Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Shaobin, Huang, Zhipeng, Zhang, Ying, Wang, Fangyikang, Fu, Canmiao, Yang, Binxin, Sun, Chong, Li, Chen, Wang, Yali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915295475531776
author Zhuang, Shaobin
Huang, Zhipeng
Zhang, Ying
Wang, Fangyikang
Fu, Canmiao
Yang, Binxin
Sun, Chong
Li, Chen
Wang, Yali
author_facet Zhuang, Shaobin
Huang, Zhipeng
Zhang, Ying
Wang, Fangyikang
Fu, Canmiao
Yang, Binxin
Sun, Chong
Li, Chen
Wang, Yali
contents GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream. The project page is at https://zhuangshaobin.github.io/Video-GPT.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12489
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-GPT via Next Clip Diffusion
Zhuang, Shaobin
Huang, Zhipeng
Zhang, Ying
Wang, Fangyikang
Fu, Canmiao
Yang, Binxin
Sun, Chong
Li, Chen
Wang, Yali
Computer Vision and Pattern Recognition
Artificial Intelligence
GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream. The project page is at https://zhuangshaobin.github.io/Video-GPT.github.io/.
title Video-GPT via Next Clip Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.12489