L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tiange, Huang, Zhimeng, Meng, Xiandong, Zhang, Kai, Deng, Zhipin, Ma, Siwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912764281225216
author Zhang, Tiange
Huang, Zhimeng
Meng, Xiandong
Zhang, Kai
Deng, Zhipin
Ma, Siwei
author_facet Zhang, Tiange
Huang, Zhimeng
Meng, Xiandong
Zhang, Kai
Deng, Zhipin
Ma, Siwei
contents Neural Video Compression has emerged in recent years, with condition-based frameworks outperforming traditional codecs. However, most existing methods rely solely on the previous frame's features to predict temporal context, leading to two critical issues. First, the short reference window misses long-term dependencies and fine texture details. Second, propagating only feature-level information accumulates errors over frames, causing prediction inaccuracies and loss of subtle textures. To address these, we propose the Long-term Spatio-Temporal Enhanced Context (L-STEC) method. We first extend the reference chain with LSTM to capture long-term dependencies. We then incorporate warped spatial context from the pixel domain, fusing spatio-temporal information through a multi-receptive field network to better preserve reference details. Experimental results show that L-STEC significantly improves compression by enriching contextual information, achieving 37.01% bitrate savings in PSNR and 31.65% in MS-SSIM compared to DCVC-TCM, outperforming both VTM-17.0 and DCVC-FM and establishing new state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
Zhang, Tiange
Huang, Zhimeng
Meng, Xiandong
Zhang, Kai
Deng, Zhipin
Ma, Siwei
Computer Vision and Pattern Recognition
I.4.2
Neural Video Compression has emerged in recent years, with condition-based frameworks outperforming traditional codecs. However, most existing methods rely solely on the previous frame's features to predict temporal context, leading to two critical issues. First, the short reference window misses long-term dependencies and fine texture details. Second, propagating only feature-level information accumulates errors over frames, causing prediction inaccuracies and loss of subtle textures. To address these, we propose the Long-term Spatio-Temporal Enhanced Context (L-STEC) method. We first extend the reference chain with LSTM to capture long-term dependencies. We then incorporate warped spatial context from the pixel domain, fusing spatio-temporal information through a multi-receptive field network to better preserve reference details. Experimental results show that L-STEC significantly improves compression by enriching contextual information, achieving 37.01% bitrate savings in PSNR and 31.65% in MS-SSIM compared to DCVC-TCM, outperforming both VTM-17.0 and DCVC-FM and establishing new state-of-the-art performance.
title L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
topic Computer Vision and Pattern Recognition
I.4.2
url https://arxiv.org/abs/2512.12790