Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Koyuncu, A. Burakhan, Gao, Han, Boev, Atanas, Gaikov, Georgii, Alshina, Elena, Steinbach, Eckehard
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909120527859712
author Koyuncu, A. Burakhan
Gao, Han
Boev, Atanas
Gaikov, Georgii
Alshina, Elena
Steinbach, Eckehard
author_facet Koyuncu, A. Burakhan
Gao, Han
Boev, Atanas
Gaikov, Georgii
Alshina, Elena
Steinbach, Eckehard
contents Entropy modeling is a key component for high-performance image compression algorithms. Recent developments in autoregressive context modeling helped learning-based methods to surpass their classical counterparts. However, the performance of those models can be further improved due to the underexploited spatio-channel dependencies in latent space, and the suboptimal implementation of context adaptivity. Inspired by the adaptive characteristics of the transformers, we propose a transformer-based context model, named Contextformer, which generalizes the de facto standard attention mechanism to spatio-channel attention. We replace the context model of a modern compression framework with the Contextformer and test it on the widely used Kodak, CLIC2020, and Tecnick image datasets. Our experimental results show that the proposed model provides up to 11% rate savings compared to the standard Versatile Video Coding (VVC) Test Model (VTM) 16.2, and outperforms various learning-based models in terms of PSNR and MS-SSIM.
format Preprint
id arxiv_https___arxiv_org_abs_2203_02452
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
Koyuncu, A. Burakhan
Gao, Han
Boev, Atanas
Gaikov, Georgii
Alshina, Elena
Steinbach, Eckehard
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Entropy modeling is a key component for high-performance image compression algorithms. Recent developments in autoregressive context modeling helped learning-based methods to surpass their classical counterparts. However, the performance of those models can be further improved due to the underexploited spatio-channel dependencies in latent space, and the suboptimal implementation of context adaptivity. Inspired by the adaptive characteristics of the transformers, we propose a transformer-based context model, named Contextformer, which generalizes the de facto standard attention mechanism to spatio-channel attention. We replace the context model of a modern compression framework with the Contextformer and test it on the widely used Kodak, CLIC2020, and Tecnick image datasets. Our experimental results show that the proposed model provides up to 11% rate savings compared to the standard Versatile Video Coding (VVC) Test Model (VTM) 16.2, and outperforms various learning-based models in terms of PSNR and MS-SSIM.
title Contextformer: A Transformer with Spatio-Channel Attention for Context Modeling in Learned Image Compression
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2203.02452