When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Pingping, Li, Jinlong, Chen, Kecheng, Wang, Meng, Xu, Long, Li, Haoliang, Sebe, Nicu, Kwong, Sam, Wang, Shiqi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915151320449024
author Zhang, Pingping
Li, Jinlong
Chen, Kecheng
Wang, Meng
Xu, Long
Li, Haoliang
Sebe, Nicu
Kwong, Sam
Wang, Shiqi
author_facet Zhang, Pingping
Li, Jinlong
Chen, Kecheng
Wang, Meng
Xu, Long
Li, Haoliang
Sebe, Nicu
Kwong, Sam
Wang, Shiqi
contents Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video compression. Herein, we introduce a unified paradigm for Cross-Modality Video Coding (CMVC), which is a pioneering approach to explore multimodality representation and video generative models in video coding. Specifically, on the encoder side, we disentangle a video into spatial content and motion components, which are subsequently transformed into distinct modalities to achieve very compact representation by leveraging MLLMs. During decoding, previously encoded components and video generation models are leveraged to create multiple encoding-decoding modes that optimize video reconstruction quality for specific decoding requirements, including Text-Text-to-Video (TT2V) mode to ensure high-quality semantic information and Image-Text-to-Video (IT2V) mode to achieve superb perceptual consistency. In addition, we propose an efficient frame interpolation model for IT2V mode via Low-Rank Adaption (LoRA) tuning to guarantee perceptual quality, which allows the generated motion cues to behave smoothly. Experiments on benchmarks indicate that TT2V achieves effective semantic reconstruction, while IT2V exhibits competitive perceptual consistency. These results highlight potential directions for future research in video coding.
format Preprint
id arxiv_https___arxiv_org_abs_2408_08093
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
Zhang, Pingping
Li, Jinlong
Chen, Kecheng
Wang, Meng
Xu, Long
Li, Haoliang
Sebe, Nicu
Kwong, Sam
Wang, Shiqi
Computer Vision and Pattern Recognition
Multimedia
Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video compression. Herein, we introduce a unified paradigm for Cross-Modality Video Coding (CMVC), which is a pioneering approach to explore multimodality representation and video generative models in video coding. Specifically, on the encoder side, we disentangle a video into spatial content and motion components, which are subsequently transformed into distinct modalities to achieve very compact representation by leveraging MLLMs. During decoding, previously encoded components and video generation models are leveraged to create multiple encoding-decoding modes that optimize video reconstruction quality for specific decoding requirements, including Text-Text-to-Video (TT2V) mode to ensure high-quality semantic information and Image-Text-to-Video (IT2V) mode to achieve superb perceptual consistency. In addition, we propose an efficient frame interpolation model for IT2V mode via Low-Rank Adaption (LoRA) tuning to guarantee perceptual quality, which allows the generated motion cues to behave smoothly. Experiments on benchmarks indicate that TT2V achieves effective semantic reconstruction, while IT2V exhibits competitive perceptual consistency. These results highlight potential directions for future research in video coding.
title When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2408.08093