DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yan, Cheng, Zhengxue, Zhang, Junxuan, Gu, Qunshan, Wang, Qi, Song, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908374289874944
author Zhao, Yan
Cheng, Zhengxue
Zhang, Junxuan
Gu, Qunshan
Wang, Qi
Song, Li
author_facet Zhao, Yan
Cheng, Zhengxue
Zhang, Junxuan
Gu, Qunshan
Wang, Qi
Song, Li
contents Most learning-based lossless compressors are designed for a single modality, requiring separate models for multi-modal data and lacking flexibility. However, different modalities vary significantly in format and statistical properties, making it ineffective to use compressors that lack modality-specific adaptations. While multi-modal large language models (MLLMs) offer a potential solution for modality-unified compression, their excessive complexity hinders practical deployment. To address these challenges, we focus on the two most common modalities, image and text, and propose DualComp, the first unified and lightweight learning-based dual-modality lossless compressor. Built on a lightweight backbone, DualComp incorporates three key structural enhancements to handle modality heterogeneity: modality-unified tokenization, modality-switching contextual learning, and modality-routing mixture-of-experts. A reparameterization training strategy is also used to boost compression performance. DualComp integrates both modality-specific and shared parameters for efficient parameter utilization, enabling near real-time inference (200KB/s) on desktop CPUs. With much fewer parameters, DualComp achieves compression performance on par with the SOTA LLM-based methods for both text and image datasets. Its simplified single-modality variant surpasses the previous best image compressor on the Kodak dataset by about 9% using just 1.2% of the model size.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16256
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
Zhao, Yan
Cheng, Zhengxue
Zhang, Junxuan
Gu, Qunshan
Wang, Qi
Song, Li
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Most learning-based lossless compressors are designed for a single modality, requiring separate models for multi-modal data and lacking flexibility. However, different modalities vary significantly in format and statistical properties, making it ineffective to use compressors that lack modality-specific adaptations. While multi-modal large language models (MLLMs) offer a potential solution for modality-unified compression, their excessive complexity hinders practical deployment. To address these challenges, we focus on the two most common modalities, image and text, and propose DualComp, the first unified and lightweight learning-based dual-modality lossless compressor. Built on a lightweight backbone, DualComp incorporates three key structural enhancements to handle modality heterogeneity: modality-unified tokenization, modality-switching contextual learning, and modality-routing mixture-of-experts. A reparameterization training strategy is also used to boost compression performance. DualComp integrates both modality-specific and shared parameters for efficient parameter utilization, enabling near real-time inference (200KB/s) on desktop CPUs. With much fewer parameters, DualComp achieves compression performance on par with the SOTA LLM-based methods for both text and image datasets. Its simplified single-modality variant surpasses the previous best image compressor on the Kodak dataset by about 9% using just 1.2% of the model size.
title DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2505.16256