Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lau, Kin Wai, Rehman, Yasar Abbas Ur, Po, Lai-Man, de Gusmão, Pedro Porto Buarque
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917316399202304
author Lau, Kin Wai
Rehman, Yasar Abbas Ur
Po, Lai-Man
de Gusmão, Pedro Porto Buarque
author_facet Lau, Kin Wai
Rehman, Yasar Abbas Ur
Po, Lai-Man
de Gusmão, Pedro Porto Buarque
contents Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Expert (MoE) architectures with specialized experts and routing, they still inflate parameter counts and introduce routing overhead. In this paper, we propose Omni-C (Omni-Compress), a single dense Transformer-based encoder that learns competitive shared representations across heterogeneous modalities--images, audio, and text--through unimodal contrastive pretraining on large-scale unaligned data. By maximizing parameter sharing in the backbone and using lightweight modality-specific projection heads, Omni-C effectively mitigates inter-modality conflicts without requiring MoE, paired supervision, or routing. This design supports efficient deployment on memory-constrained systems via sequential modality processing and low-memory inference, eliminating the need for parallel expert loading or specialized hardware. Experiments show Omni-C achieves performance comparable to expert models in unimodal and cross-model tasks, with modest zero-shot degradation on audio and text that is largely recovered through lightweight linear probing or parameter efficient fine-tuning. The unified architecture substantially reduces inference memory usage compared to multi-encoder baselines, advancing efficient and scalable multimodal learning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05528
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
Lau, Kin Wai
Rehman, Yasar Abbas Ur
Po, Lai-Man
de Gusmão, Pedro Porto Buarque
Multimedia
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Expert (MoE) architectures with specialized experts and routing, they still inflate parameter counts and introduce routing overhead. In this paper, we propose Omni-C (Omni-Compress), a single dense Transformer-based encoder that learns competitive shared representations across heterogeneous modalities--images, audio, and text--through unimodal contrastive pretraining on large-scale unaligned data. By maximizing parameter sharing in the backbone and using lightweight modality-specific projection heads, Omni-C effectively mitigates inter-modality conflicts without requiring MoE, paired supervision, or routing. This design supports efficient deployment on memory-constrained systems via sequential modality processing and low-memory inference, eliminating the need for parallel expert loading or specialized hardware. Experiments show Omni-C achieves performance comparable to expert models in unimodal and cross-model tasks, with modest zero-shot degradation on audio and text that is largely recovered through lightweight linear probing or parameter efficient fine-tuning. The unified architecture substantially reduces inference memory usage compared to multi-encoder baselines, advancing efficient and scalable multimodal learning.
title Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
topic Multimedia
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2603.05528