Semantic Compression via Multimodal Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Grassucci, Eleonora, Cicchetti, Giordano, Uncini, Aurelio, Comminiello, Danilo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183577022464
author Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
author_facet Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
contents Multimodal representation learning produces high-dimensional embeddings that align diverse modalities in a shared latent space. While this enables strong generalization, it also introduces scalability challenges, both in terms of storage and downstream processing. A key open problem is how to achieve semantic compression, reducing the memory footprint of multimodal embeddings while preserving their ability to represent shared semantic content across modalities. In this paper, we prove a strong connection between reducing the modality gap, which is the residual separation of embeddings from different modalities, and the feasibility of post-training semantic compression. When the gap is sufficiently reduced, embeddings from different modalities but expressing the same semantics share a common portion of the space. Therefore, their centroid is a faithful representation of such a semantic concept. This enables replacing multiple embeddings with a single centroid, yielding significant memory savings. We propose a novel approach for semantic compression grounded on the latter intuition, operating directly on pretrained encoders. We demonstrate its effectiveness across diverse large-scale multimodal downstream tasks. Our results highlight that modality alignment is a key enabler for semantic compression, showing that the proposed approach achieves significant compression without sacrificing performance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24431
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic Compression via Multimodal Representation Learning
Grassucci, Eleonora
Cicchetti, Giordano
Uncini, Aurelio
Comminiello, Danilo
Machine Learning
Multimodal representation learning produces high-dimensional embeddings that align diverse modalities in a shared latent space. While this enables strong generalization, it also introduces scalability challenges, both in terms of storage and downstream processing. A key open problem is how to achieve semantic compression, reducing the memory footprint of multimodal embeddings while preserving their ability to represent shared semantic content across modalities. In this paper, we prove a strong connection between reducing the modality gap, which is the residual separation of embeddings from different modalities, and the feasibility of post-training semantic compression. When the gap is sufficiently reduced, embeddings from different modalities but expressing the same semantics share a common portion of the space. Therefore, their centroid is a faithful representation of such a semantic concept. This enables replacing multiple embeddings with a single centroid, yielding significant memory savings. We propose a novel approach for semantic compression grounded on the latter intuition, operating directly on pretrained encoders. We demonstrate its effectiveness across diverse large-scale multimodal downstream tasks. Our results highlight that modality alignment is a key enabler for semantic compression, showing that the proposed approach achieves significant compression without sacrificing performance.
title Semantic Compression via Multimodal Representation Learning
topic Machine Learning
url https://arxiv.org/abs/2509.24431