Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Zhening, Fox, Gereon, Zeng, Yanhong, Pan, Xingang, Elgharib, Mohamed, Theobalt, Christian, Chen, Kai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913427249692672
author Xing, Zhening
Fox, Gereon
Zeng, Yanhong
Pan, Xingang
Elgharib, Mohamed
Theobalt, Christian
Chen, Kai
author_facet Xing, Zhening
Fox, Gereon
Zeng, Yanhong
Pan, Xingang
Elgharib, Mohamed
Theobalt, Christian
Chen, Kai
contents Large Language Models have shown remarkable efficacy in generating streaming data such as text and audio, thanks to their temporally uni-directional attention mechanism, which models correlations between the current token and previous tokens. However, video streaming remains much less explored, despite a growing need for live video processing. State-of-the-art video diffusion models leverage bi-directional temporal attention to model the correlations between the current frame and all the surrounding (i.e. including future) frames, which hinders them from processing streaming videos. To address this problem, we present Live2Diff, the first attempt at designing a video diffusion model with uni-directional temporal attention, specifically targeting live streaming video translation. Compared to previous works, our approach ensures temporal consistency and smoothness by correlating the current frame with its predecessors and a few initial warmup frames, without any future frames. Additionally, we use a highly efficient denoising scheme featuring a KV-cache mechanism and pipelining, to facilitate streaming video translation at interactive framerates. Extensive experiments demonstrate the effectiveness of the proposed attention mechanism and pipeline, outperforming previous methods in terms of temporal smoothness and/or efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08701
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
Xing, Zhening
Fox, Gereon
Zeng, Yanhong
Pan, Xingang
Elgharib, Mohamed
Theobalt, Christian
Chen, Kai
Computer Vision and Pattern Recognition
Large Language Models have shown remarkable efficacy in generating streaming data such as text and audio, thanks to their temporally uni-directional attention mechanism, which models correlations between the current token and previous tokens. However, video streaming remains much less explored, despite a growing need for live video processing. State-of-the-art video diffusion models leverage bi-directional temporal attention to model the correlations between the current frame and all the surrounding (i.e. including future) frames, which hinders them from processing streaming videos. To address this problem, we present Live2Diff, the first attempt at designing a video diffusion model with uni-directional temporal attention, specifically targeting live streaming video translation. Compared to previous works, our approach ensures temporal consistency and smoothness by correlating the current frame with its predecessors and a few initial warmup frames, without any future frames. Additionally, we use a highly efficient denoising scheme featuring a KV-cache mechanism and pipelining, to facilitate streaming video translation at interactive framerates. Extensive experiments demonstrate the effectiveness of the proposed attention mechanism and pipeline, outperforming previous methods in terms of temporal smoothness and/or efficiency.
title Live2Diff: Live Stream Translation via Uni-directional Attention in Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.08701