Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xu, Lu, Ming, Chen, Yan, Ma, Zhan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908430357233664
author Zhang, Xu
Lu, Ming
Chen, Yan
Ma, Zhan
author_facet Zhang, Xu
Lu, Ming
Chen, Yan
Ma, Zhan
contents In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization tends to yield latent spaces with limited semantic richness, which hinders effective semantic inference in downstream tasks. Moreover, achieving high performance with these models often requires fine-tuning the entire vision model, which is computationally intensive, especially for large models. To address these problems, we introduce Perception-Oriented Latent Coding (POLC), an approach that enriches the semantic content of latent features for high-performance compressed domain semantic inference. With the semantically rich latent space, POLC requires only a plug-and-play adapter for fine-tuning, significantly reducing the parameter count compared to previous MSE-oriented methods. Experimental results demonstrate that POLC achieves rate-perception performance comparable to state-of-the-art generative image coding methods while markedly enhancing performance in vision tasks, with minimal fine-tuning overhead. Code is available at https://github.com/NJUVISION/POLC.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
Zhang, Xu
Lu, Ming
Chen, Yan
Ma, Zhan
Computer Vision and Pattern Recognition
Image and Video Processing
In recent years, compressed domain semantic inference has primarily relied on learned image coding models optimized for mean squared error (MSE). However, MSE-oriented optimization tends to yield latent spaces with limited semantic richness, which hinders effective semantic inference in downstream tasks. Moreover, achieving high performance with these models often requires fine-tuning the entire vision model, which is computationally intensive, especially for large models. To address these problems, we introduce Perception-Oriented Latent Coding (POLC), an approach that enriches the semantic content of latent features for high-performance compressed domain semantic inference. With the semantically rich latent space, POLC requires only a plug-and-play adapter for fine-tuning, significantly reducing the parameter count compared to previous MSE-oriented methods. Experimental results demonstrate that POLC achieves rate-perception performance comparable to state-of-the-art generative image coding methods while markedly enhancing performance in vision tasks, with minimal fine-tuning overhead. Code is available at https://github.com/NJUVISION/POLC.
title Perception-Oriented Latent Coding for High-Performance Compressed Domain Semantic Inference
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2507.01608