OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianchao, Yu, Shujian, Zu, Xinrui, Wei, Zhaolong, Gummeson, Jeremy, Cheng, Jack C. P., Jenssen, Robert
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914614225141760
author Li, Tianchao
Yu, Shujian
Zu, Xinrui
Wei, Zhaolong
Gummeson, Jeremy
Cheng, Jack C. P.
Jenssen, Robert
author_facet Li, Tianchao
Yu, Shujian
Zu, Xinrui
Wei, Zhaolong
Gummeson, Jeremy
Cheng, Jack C. P.
Jenssen, Robert
contents Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent two-way comparisons and therefore do not explicitly model higher-order dependencies among multiple modalities. Recent beyond-pairwise objectives approach this problem from statistical or geometric perspectives, but arbitrary-modality alignment still lacks a principled criterion for defining what each modality should preserve and compress relative to the others. We revisit arbitrary-modality alignment through the Information Bottleneck principle. In multi-modal learning, sufficiency should preserve information predictable from the remaining modalities, while minimality should compress modality-specific information not supported by them. This naturally leads to a One-vs-All view, where each modality is characterized with respect to the remaining modalities. We propose OVA-IB, an Information Bottleneck framework for arbitrary-modality alignment. OVA-IB optimizes a tractable One-vs-All contrastive lower bound for sufficiency connected to a Dual Total Correlation-style objective, uses a parameter-free geometry-aware projection score, and derives a tractable upper-bound regularizer for minimality by bounding each representation's dependence on its own input with representation distributions induced by the remaining modalities. Experiments on classification, regression, modality-agnostic evaluation, and cross-modal retrieval benchmarks demonstrate strong and robust performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment
Li, Tianchao
Yu, Shujian
Zu, Xinrui
Wei, Zhaolong
Gummeson, Jeremy
Cheng, Jack C. P.
Jenssen, Robert
Machine Learning
Information Theory
Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent two-way comparisons and therefore do not explicitly model higher-order dependencies among multiple modalities. Recent beyond-pairwise objectives approach this problem from statistical or geometric perspectives, but arbitrary-modality alignment still lacks a principled criterion for defining what each modality should preserve and compress relative to the others. We revisit arbitrary-modality alignment through the Information Bottleneck principle. In multi-modal learning, sufficiency should preserve information predictable from the remaining modalities, while minimality should compress modality-specific information not supported by them. This naturally leads to a One-vs-All view, where each modality is characterized with respect to the remaining modalities. We propose OVA-IB, an Information Bottleneck framework for arbitrary-modality alignment. OVA-IB optimizes a tractable One-vs-All contrastive lower bound for sufficiency connected to a Dual Total Correlation-style objective, uses a parameter-free geometry-aware projection score, and derives a tractable upper-bound regularizer for minimality by bounding each representation's dependence on its own input with representation distributions induced by the remaining modalities. Experiments on classification, regression, modality-agnostic evaluation, and cross-modal retrieval benchmarks demonstrate strong and robust performance.
title OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2605.29900