StreamDiT: Real-Time Streaming Text-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kodaira, Akio, Hou, Tingbo, Hou, Ji, Georgopoulos, Markos, Juefei-Xu, Felix, Tomizuka, Masayoshi, Zhao, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908916498038784
author Kodaira, Akio
Hou, Tingbo
Hou, Ji
Georgopoulos, Markos
Juefei-Xu, Felix
Tomizuka, Masayoshi
Zhao, Yue
author_facet Kodaira, Akio
Hou, Tingbo
Hou, Ji
Georgopoulos, Markos
Juefei-Xu, Felix
Tomizuka, Masayoshi
Zhao, Yue
contents Recently, great progress has been achieved in text-to-video (T2V) generation by scaling transformer-based diffusion models to billions of parameters, which can generate high-quality videos. However, existing models typically produce only short clips offline, restricting their use cases in interactive and real-time applications. This paper addresses these challenges by proposing StreamDiT, a streaming video generation model. StreamDiT training is based on flow matching by adding a moving buffer. We design mixed training with different partitioning schemes of buffered frames to boost both content consistency and visual quality. StreamDiT modeling is based on adaLN DiT with varying time embedding and window attention. To practice the proposed method, we train a StreamDiT model with 4B parameters. In addition, we propose a multistep distillation method tailored for StreamDiT. Sampling distillation is performed in each segment of a chosen partitioning scheme. After distillation, the total number of function evaluations (NFEs) is reduced to the number of chunks in a buffer. Finally, our distilled model reaches real-time performance at 16 FPS on one GPU, which can generate video streams at 512p resolution. We evaluate our method through both quantitative metrics and human evaluation. Our model enables real-time applications, e.g. streaming generation, interactive generation, and video-to-video. We provide video results and more examples in our project website: https://cumulo-autumn.github.io/StreamDiT/
format Preprint
id arxiv_https___arxiv_org_abs_2507_03745
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StreamDiT: Real-Time Streaming Text-to-Video Generation
Kodaira, Akio
Hou, Tingbo
Hou, Ji
Georgopoulos, Markos
Juefei-Xu, Felix
Tomizuka, Masayoshi
Zhao, Yue
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
Recently, great progress has been achieved in text-to-video (T2V) generation by scaling transformer-based diffusion models to billions of parameters, which can generate high-quality videos. However, existing models typically produce only short clips offline, restricting their use cases in interactive and real-time applications. This paper addresses these challenges by proposing StreamDiT, a streaming video generation model. StreamDiT training is based on flow matching by adding a moving buffer. We design mixed training with different partitioning schemes of buffered frames to boost both content consistency and visual quality. StreamDiT modeling is based on adaLN DiT with varying time embedding and window attention. To practice the proposed method, we train a StreamDiT model with 4B parameters. In addition, we propose a multistep distillation method tailored for StreamDiT. Sampling distillation is performed in each segment of a chosen partitioning scheme. After distillation, the total number of function evaluations (NFEs) is reduced to the number of chunks in a buffer. Finally, our distilled model reaches real-time performance at 16 FPS on one GPU, which can generate video streams at 512p resolution. We evaluate our method through both quantitative metrics and human evaluation. Our model enables real-time applications, e.g. streaming generation, interactive generation, and video-to-video. We provide video results and more examples in our project website: https://cumulo-autumn.github.io/StreamDiT/
title StreamDiT: Real-Time Streaming Text-to-Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2507.03745