ContentV: Efficient Training of Video Generation Models with Limited Compute

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Wenfeng, Chen, Renjie, Liu, Boyuan, Yan, Shiyue, Feng, Ruoyu, Wei, Jiangchuan, Zhang, Yichen, Zhou, Yimeng, Feng, Chao, Ran, Jiao, Wu, Qi, Liu, Zuotao, Guo, Mingyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910999832952832
author Lin, Wenfeng
Chen, Renjie
Liu, Boyuan
Yan, Shiyue
Feng, Ruoyu
Wei, Jiangchuan
Zhang, Yichen
Zhou, Yimeng
Feng, Chao
Ran, Jiao
Wu, Qi
Liu, Zuotao
Guo, Mingyu
author_facet Lin, Wenfeng
Chen, Renjie
Liu, Boyuan
Yan, Shiyue
Feng, Ruoyu
Wei, Jiangchuan
Zhang, Yichen
Zhou, Yimeng
Feng, Chao
Ran, Jiao
Wu, Qi
Liu, Zuotao
Guo, Mingyu
contents Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter text-to-video model that achieves state-of-the-art performance (85.14 on VBench) after training on 256 x 64GB Neural Processing Units (NPUs) for merely four weeks. ContentV generates diverse, high-quality videos across multiple resolutions and durations from text prompts, enabled by three key innovations: (1) A minimalist architecture that maximizes reuse of pre-trained image generation models for video generation; (2) A systematic multi-stage training strategy leveraging flow matching for enhanced efficiency; and (3) A cost-effective reinforcement learning with human feedback framework that improves generation quality without requiring additional human annotations. All the code and models are available at: https://contentv.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ContentV: Efficient Training of Video Generation Models with Limited Compute
Lin, Wenfeng
Chen, Renjie
Liu, Boyuan
Yan, Shiyue
Feng, Ruoyu
Wei, Jiangchuan
Zhang, Yichen
Zhou, Yimeng
Feng, Chao
Ran, Jiao
Wu, Qi
Liu, Zuotao
Guo, Mingyu
Computer Vision and Pattern Recognition
Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter text-to-video model that achieves state-of-the-art performance (85.14 on VBench) after training on 256 x 64GB Neural Processing Units (NPUs) for merely four weeks. ContentV generates diverse, high-quality videos across multiple resolutions and durations from text prompts, enabled by three key innovations: (1) A minimalist architecture that maximizes reuse of pre-trained image generation models for video generation; (2) A systematic multi-stage training strategy leveraging flow matching for enhanced efficiency; and (3) A cost-effective reinforcement learning with human feedback framework that improves generation quality without requiring additional human annotations. All the code and models are available at: https://contentv.github.io.
title ContentV: Efficient Training of Video Generation Models with Limited Compute
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05343