Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Qifan, Zou, Jiale, Zhang, Jinhua, Long, Wei, Zhou, Xingyu, Gu, Shuhang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917355860262912
author Li, Qifan
Zou, Jiale
Zhang, Jinhua
Long, Wei
Zhou, Xingyu
Gu, Shuhang
author_facet Li, Qifan
Zou, Jiale
Zhang, Jinhua
Long, Wei
Zhou, Xingyu
Gu, Shuhang
contents Vector-quantized based models have recently demonstrated strong potential for visual prior modeling. However, existing VQ-based methods simply encode visual features with nearest codebook items and train index predictor with code-level supervision. Due to the richness of visual signal, VQ encoding often leads to large quantization error. Furthermore, training predictor with code-level supervision can not take the final reconstruction errors into consideration, result in sub-optimal prior modeling accuracy. In this paper we address the above two issues and propose a Texture Vector-Quantization and a Reconstruction Aware Prediction strategy. The texture vector-quantization strategy leverages the task character of super-resolution and only introduce codebook to model the prior of missing textures. While the reconstruction aware prediction strategy makes use of the straight-through estimator to directly train index predictor with image-level supervision. Our proposed generative SR model (TVQ&RAP) is able to deliver photo-realistic SR results with small computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
Li, Qifan
Zou, Jiale
Zhang, Jinhua
Long, Wei
Zhou, Xingyu
Gu, Shuhang
Computer Vision and Pattern Recognition
Vector-quantized based models have recently demonstrated strong potential for visual prior modeling. However, existing VQ-based methods simply encode visual features with nearest codebook items and train index predictor with code-level supervision. Due to the richness of visual signal, VQ encoding often leads to large quantization error. Furthermore, training predictor with code-level supervision can not take the final reconstruction errors into consideration, result in sub-optimal prior modeling accuracy. In this paper we address the above two issues and propose a Texture Vector-Quantization and a Reconstruction Aware Prediction strategy. The texture vector-quantization strategy leverages the task character of super-resolution and only introduce codebook to model the prior of missing textures. While the reconstruction aware prediction strategy makes use of the straight-through estimator to directly train index predictor with image-level supervision. Our proposed generative SR model (TVQ&RAP) is able to deliver photo-realistic SR results with small computational cost.
title Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.23774