LongCat-Video Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meituan LongCat Team, Cai, Xunliang, Huang, Qilong, Kang, Zhuoliang, Li, Hongyu, Liang, Shijun, Ma, Liya, Ren, Siyu, Wei, Xiaoming, Xie, Rixu, Zhang, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915581583687680
author Meituan LongCat Team
Cai, Xunliang
Huang, Qilong
Kang, Zhuoliang
Li, Hongyu
Liang, Shijun
Ma, Liya
Ren, Siyu
Wei, Xiaoming
Xie, Rixu
Zhang, Tong
author_facet Meituan LongCat Team
Cai, Xunliang
Huang, Qilong
Kang, Zhuoliang
Li, Hongyu
Liang, Shijun
Ma, Liya
Ren, Siyu
Wei, Xiaoming
Xie, Rixu
Zhang, Tong
contents Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameters, delivering strong performance across multiple video generation tasks. It particularly excels in efficient and high-quality long video generation, representing our first step toward world models. Key features include: Unified architecture for multiple tasks: Built on the Diffusion Transformer (DiT) framework, LongCat-Video supports Text-to-Video, Image-to-Video, and Video-Continuation tasks with a single model; Long video generation: Pretraining on Video-Continuation tasks enables LongCat-Video to maintain high quality and temporal coherence in the generation of minutes-long videos; Efficient inference: LongCat-Video generates 720p, 30fps videos within minutes by employing a coarse-to-fine generation strategy along both the temporal and spatial axes. Block Sparse Attention further enhances efficiency, particularly at high resolutions; Strong performance with multi-reward RLHF: Multi-reward RLHF training enables LongCat-Video to achieve performance on par with the latest closed-source and leading open-source models. Code and model weights are publicly available to accelerate progress in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22200
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongCat-Video Technical Report
Meituan LongCat Team
Cai, Xunliang
Huang, Qilong
Kang, Zhuoliang
Li, Hongyu
Liang, Shijun
Ma, Liya
Ren, Siyu
Wei, Xiaoming
Xie, Rixu
Zhang, Tong
Computer Vision and Pattern Recognition
Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameters, delivering strong performance across multiple video generation tasks. It particularly excels in efficient and high-quality long video generation, representing our first step toward world models. Key features include: Unified architecture for multiple tasks: Built on the Diffusion Transformer (DiT) framework, LongCat-Video supports Text-to-Video, Image-to-Video, and Video-Continuation tasks with a single model; Long video generation: Pretraining on Video-Continuation tasks enables LongCat-Video to maintain high quality and temporal coherence in the generation of minutes-long videos; Efficient inference: LongCat-Video generates 720p, 30fps videos within minutes by employing a coarse-to-fine generation strategy along both the temporal and spatial axes. Block Sparse Attention further enhances efficiency, particularly at high resolutions; Strong performance with multi-reward RLHF: Multi-reward RLHF training enables LongCat-Video to achieve performance on par with the latest closed-source and leading open-source models. Code and model weights are publicly available to accelerate progress in the field.
title LongCat-Video Technical Report
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22200