DeRA: Decoupled Representation Alignment for Video Tokenization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Pengbo, Wang, Junke, Xing, Zhen, Liu, Chengxu, Dong, Daoguo, Qian, Xueming, Wu, Zuxuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917124502454272
author Guo, Pengbo
Wang, Junke
Xing, Zhen
Liu, Chengxu
Dong, Daoguo
Qian, Xueming
Wu, Zuxuan
author_facet Guo, Pengbo
Wang, Junke
Xing, Zhen
Liu, Chengxu
Dong, Daoguo
Qian, Xueming
Wu, Zuxuan
contents This paper presents DeRA, a novel 1D video tokenizer that decouples the spatial-temporal representation learning in video tokenization to achieve better training efficiency and performance. Specifically, DeRA maintains a compact 1D latent space while factorizing video encoding into appearance and motion streams, which are aligned with pretrained vision foundation models to capture the spatial semantics and temporal dynamics in videos separately. To address the gradient conflicts introduced by the heterogeneous supervision, we further propose the Symmetric Alignment-Conflict Projection (SACP) module that proactively reformulates gradients by suppressing the components along conflicting directions. Extensive experiments demonstrate that DeRA outperforms LARP, the previous state-of-the-art video tokenizer by 25% on UCF-101 in terms of rFVD. Moreover, using DeRA for autoregressive video generation, we also achieve new state-of-the-art results on both UCF-101 class-conditional generation and K600 frame prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeRA: Decoupled Representation Alignment for Video Tokenization
Guo, Pengbo
Wang, Junke
Xing, Zhen
Liu, Chengxu
Dong, Daoguo
Qian, Xueming
Wu, Zuxuan
Computer Vision and Pattern Recognition
This paper presents DeRA, a novel 1D video tokenizer that decouples the spatial-temporal representation learning in video tokenization to achieve better training efficiency and performance. Specifically, DeRA maintains a compact 1D latent space while factorizing video encoding into appearance and motion streams, which are aligned with pretrained vision foundation models to capture the spatial semantics and temporal dynamics in videos separately. To address the gradient conflicts introduced by the heterogeneous supervision, we further propose the Symmetric Alignment-Conflict Projection (SACP) module that proactively reformulates gradients by suppressing the components along conflicting directions. Extensive experiments demonstrate that DeRA outperforms LARP, the previous state-of-the-art video tokenizer by 25% on UCF-101 in terms of rFVD. Moreover, using DeRA for autoregressive video generation, we also achieve new state-of-the-art results on both UCF-101 class-conditional generation and K600 frame prediction.
title DeRA: Decoupled Representation Alignment for Video Tokenization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04483