End-to-End Semantic Preservation in Text-Aware Image Compression Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Della Fiore, Stefano, Gnutti, Alessandro, Dalai, Marco, Migliorati, Pierangelo, Leonardi, Riccardo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908593121394688
author Della Fiore, Stefano
Gnutti, Alessandro
Dalai, Marco
Migliorati, Pierangelo
Leonardi, Riccardo
author_facet Della Fiore, Stefano
Gnutti, Alessandro
Dalai, Marco
Migliorati, Pierangelo
Leonardi, Riccardo
contents Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. We further extend this study to general-purpose encoders, exploring their capacity to preserve hidden semantics under extreme compression. Instead of optimizing for visual fidelity, we examine whether compact, visually degraded representations can retain recoverable meaning through learned enhancement and recognition modules. Results demonstrate that semantic information can persist despite severe compression, bridging text-oriented compression and general-purpose semantic preservation in machine-centered image coding.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-End Semantic Preservation in Text-Aware Image Compression Systems
Della Fiore, Stefano
Gnutti, Alessandro
Dalai, Marco
Migliorati, Pierangelo
Leonardi, Riccardo
Image and Video Processing
Computer Vision and Pattern Recognition
Traditional image compression methods aim to reconstruct images for human perception, prioritizing visual fidelity over task relevance. In contrast, Coding for Machines focuses on preserving information essential for automated understanding. Building on this principle, we present an end-to-end compression framework that retains text-specific features for Optical Character Recognition (OCR). The encoder operates at roughly half the computational cost of the OCR module, making it suitable for resource-limited devices. When on-device OCR is infeasible, images can be efficiently compressed and later decoded to recover textual content. Experiments show significant improvements in text extraction accuracy at low bitrates, even outperforming OCR on uncompressed images. We further extend this study to general-purpose encoders, exploring their capacity to preserve hidden semantics under extreme compression. Instead of optimizing for visual fidelity, we examine whether compact, visually degraded representations can retain recoverable meaning through learned enhancement and recognition modules. Results demonstrate that semantic information can persist despite severe compression, bridging text-oriented compression and general-purpose semantic preservation in machine-centered image coding.
title End-to-End Semantic Preservation in Text-Aware Image Compression Systems
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.19495