Augmented Deep Contexts for Spatially Embedded Video Coding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bian, Yifan, Tang, Chuanbo, Li, Li, Liu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908354153021440
author Bian, Yifan
Tang, Chuanbo
Li, Li
Liu, Dong
author_facet Bian, Yifan
Tang, Chuanbo
Li, Li
Liu, Dong
contents Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Augmented Deep Contexts for Spatially Embedded Video Coding
Bian, Yifan
Tang, Chuanbo
Li, Li
Liu, Dong
Image and Video Processing
Computer Vision and Pattern Recognition
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-of-the-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https://github.com/EsakaK/SEVC.
title Augmented Deep Contexts for Spatially Embedded Video Coding
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.05309