Exploring Compressed Image Representation as a Perceptual Proxy: A Study

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Chen-Hsiu, Wu, Ja-Ling
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913196287197184
author Huang, Chen-Hsiu
Wu, Ja-Ling
author_facet Huang, Chen-Hsiu
Wu, Ja-Ling
contents We propose an end-to-end learned image compression codec wherein the analysis transform is jointly trained with an object classification task. This study affirms that the compressed latent representation can predict human perceptual distance judgments with an accuracy comparable to a custom-tailored DNN-based quality metric. We further investigate various neural encoders and demonstrate the effectiveness of employing the analysis transform as a perceptual loss network for image tasks beyond quality judgments. Our experiments show that the off-the-shelf neural encoder proves proficient in perceptual modeling without needing an additional VGG network. We expect this research to serve as a valuable reference developing of a semantic-aware and coding-efficient neural encoder.
format Preprint
id arxiv_https___arxiv_org_abs_2401_07200
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Compressed Image Representation as a Perceptual Proxy: A Study
Huang, Chen-Hsiu
Wu, Ja-Ling
Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
We propose an end-to-end learned image compression codec wherein the analysis transform is jointly trained with an object classification task. This study affirms that the compressed latent representation can predict human perceptual distance judgments with an accuracy comparable to a custom-tailored DNN-based quality metric. We further investigate various neural encoders and demonstrate the effectiveness of employing the analysis transform as a perceptual loss network for image tasks beyond quality judgments. Our experiments show that the off-the-shelf neural encoder proves proficient in perceptual modeling without needing an additional VGG network. We expect this research to serve as a valuable reference developing of a semantic-aware and coding-efficient neural encoder.
title Exploring Compressed Image Representation as a Perceptual Proxy: A Study
topic Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2401.07200