Matten: Video Generation with Mamba-Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914789835407360 |
|---|---|
| author | Gao, Yu Huang, Jiancheng Sun, Xiaopeng Jie, Zequn Zhong, Yujie Ma, Lin |
| author_facet | Gao, Yu Huang, Jiancheng Sun, Xiaopeng Jie, Zequn Zhong, Yujie Ma, Lin |
| contents | In this paper, we introduce Matten, a cutting-edge latent diffusion model with Mamba-Attention architecture for video generation. With minimal computational cost, Matten employs spatial-temporal attention for local video content modeling and bidirectional Mamba for global video content modeling. Our comprehensive experimental evaluation demonstrates that Matten has competitive performance with the current Transformer-based and GAN-based models in benchmark performance, achieving superior FVD scores and efficiency. Additionally, we observe a direct positive correlation between the complexity of our designed model and the improvement in video quality, indicating the excellent scalability of Matten. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_03025 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Matten: Video Generation with Mamba-Attention Gao, Yu Huang, Jiancheng Sun, Xiaopeng Jie, Zequn Zhong, Yujie Ma, Lin Computer Vision and Pattern Recognition In this paper, we introduce Matten, a cutting-edge latent diffusion model with Mamba-Attention architecture for video generation. With minimal computational cost, Matten employs spatial-temporal attention for local video content modeling and bidirectional Mamba for global video content modeling. Our comprehensive experimental evaluation demonstrates that Matten has competitive performance with the current Transformer-based and GAN-based models in benchmark performance, achieving superior FVD scores and efficiency. Additionally, we observe a direct positive correlation between the complexity of our designed model and the improvement in video quality, indicating the excellent scalability of Matten. |
| title | Matten: Video Generation with Mamba-Attention |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.03025 |