RepVideo: Rethinking Cross-Layer Representation for Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Chenyang, Fan, Weichen, Lv, Zhengyao, Huang, Ziqi, Qiao, Yu, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929676996313088
author Si, Chenyang
Fan, Weichen
Lv, Zhengyao
Huang, Ziqi
Qiao, Yu
Liu, Ziwei
author_facet Si, Chenyang
Fan, Weichen
Lv, Zhengyao
Huang, Ziqi
Qiao, Yu
Liu, Ziwei
contents Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training, while offering limited insights into the direct impact of representations on the video generation process. In this paper, we initially investigate the characteristics of features in intermediate layers, finding substantial variations in attention maps across different layers. These variations lead to unstable semantic representations and contribute to cumulative differences between features, which ultimately reduce the similarity between adjacent frames and negatively affect temporal coherence. To address this, we propose RepVideo, an enhanced representation framework for text-to-video diffusion models. By accumulating features from neighboring layers to form enriched representations, this approach captures more stable semantic information. These enhanced representations are then used as inputs to the attention mechanism, thereby improving semantic expressiveness while ensuring feature consistency across adjacent frames. Extensive experiments demonstrate that our RepVideo not only significantly enhances the ability to generate accurate spatial appearances, such as capturing complex spatial relationships between multiple objects, but also improves temporal consistency in video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RepVideo: Rethinking Cross-Layer Representation for Video Generation
Si, Chenyang
Fan, Weichen
Lv, Zhengyao
Huang, Ziqi
Qiao, Yu
Liu, Ziwei
Computer Vision and Pattern Recognition
Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training, while offering limited insights into the direct impact of representations on the video generation process. In this paper, we initially investigate the characteristics of features in intermediate layers, finding substantial variations in attention maps across different layers. These variations lead to unstable semantic representations and contribute to cumulative differences between features, which ultimately reduce the similarity between adjacent frames and negatively affect temporal coherence. To address this, we propose RepVideo, an enhanced representation framework for text-to-video diffusion models. By accumulating features from neighboring layers to form enriched representations, this approach captures more stable semantic information. These enhanced representations are then used as inputs to the attention mechanism, thereby improving semantic expressiveness while ensuring feature consistency across adjacent frames. Extensive experiments demonstrate that our RepVideo not only significantly enhances the ability to generate accurate spatial appearances, such as capturing complex spatial relationships between multiple objects, but also improves temporal consistency in video generation.
title RepVideo: Rethinking Cross-Layer Representation for Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.08994