NVC-1B: A Large Neural Video Coding Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheng, Xihua, Tang, Chuanbo, Li, Li, Liu, Dong, Wu, Feng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909272129929216
author Sheng, Xihua
Tang, Chuanbo
Li, Li
Liu, Dong
Wu, Feng
author_facet Sheng, Xihua
Tang, Chuanbo
Li, Li
Liu, Dong
Wu, Feng
contents The emerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore to use different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model with more than 1 billion parameters -- NVC-1B. Experimental results show that our proposed large model achieves a significant video compression performance improvement over the small baseline model, and represents the state-of-the-art compression efficiency. We anticipate large models may bring up the video coding technologies to the next level.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19402
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NVC-1B: A Large Neural Video Coding Model
Sheng, Xihua
Tang, Chuanbo
Li, Li
Liu, Dong
Wu, Feng
Computer Vision and Pattern Recognition
Image and Video Processing
The emerging large models have achieved notable progress in the fields of natural language processing and computer vision. However, large models for neural video coding are still unexplored. In this paper, we try to explore how to build a large neural video coding model. Based on a small baseline model, we gradually scale up the model sizes of its different coding parts, including the motion encoder-decoder, motion entropy model, contextual encoder-decoder, contextual entropy model, and temporal context mining module, and analyze the influence of model sizes on video compression performance. Then, we explore to use different architectures, including CNN, mixed CNN-Transformer, and Transformer architectures, to implement the neural video coding model and analyze the influence of model architectures on video compression performance. Based on our exploration results, we design the first neural video coding model with more than 1 billion parameters -- NVC-1B. Experimental results show that our proposed large model achieves a significant video compression performance improvement over the small baseline model, and represents the state-of-the-art compression efficiency. We anticipate large models may bring up the video coding technologies to the next level.
title NVC-1B: A Large Neural Video Coding Model
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2407.19402