Allegro: Open the Black Box of Commercial-Level Video Generation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuan, Wang, Qiuyue, Cai, Yuxuan, Yang, Huan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910657974108160
author Zhou, Yuan
Wang, Qiuyue
Cai, Yuxuan
Yang, Huan
author_facet Zhou, Yuan
Wang, Qiuyue
Cai, Yuxuan
Yang, Huan
contents Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-quality models. However, despite these efforts, the available information and resources remain insufficient for achieving commercial-level performance. In this report, we open the black box and introduce $\textbf{Allegro}$, an advanced video generation model that excels in both quality and temporal consistency. We also highlight the current limitations in the field and present a comprehensive methodology for training high-performance, commercial-level video generation models, addressing key aspects such as data, model architecture, training pipeline, and evaluation. Our user study shows that Allegro surpasses existing open-source models and most commercial models, ranking just behind Hailuo and Kling. Code: https://github.com/rhymes-ai/Allegro , Model: https://huggingface.co/rhymes-ai/Allegro , Gallery: https://rhymes.ai/allegro_gallery .
format Preprint
id arxiv_https___arxiv_org_abs_2410_15458
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Allegro: Open the Black Box of Commercial-Level Video Generation Model
Zhou, Yuan
Wang, Qiuyue
Cai, Yuxuan
Yang, Huan
Computer Vision and Pattern Recognition
Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-quality models. However, despite these efforts, the available information and resources remain insufficient for achieving commercial-level performance. In this report, we open the black box and introduce $\textbf{Allegro}$, an advanced video generation model that excels in both quality and temporal consistency. We also highlight the current limitations in the field and present a comprehensive methodology for training high-performance, commercial-level video generation models, addressing key aspects such as data, model architecture, training pipeline, and evaluation. Our user study shows that Allegro surpasses existing open-source models and most commercial models, ranking just behind Hailuo and Kling. Code: https://github.com/rhymes-ai/Allegro , Model: https://huggingface.co/rhymes-ai/Allegro , Gallery: https://rhymes.ai/allegro_gallery .
title Allegro: Open the Black Box of Commercial-Level Video Generation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.15458