The Indra Representation Hypothesis for Multimodal Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Jianglin, Wang, Hailing, Yang, Kuo, Zhang, Yitian, Jenni, Simon, Fu, Yun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911569212866560
author Lu, Jianglin
Wang, Hailing
Yang, Kuo
Zhang, Yitian
Jenni, Simon
Fu, Yun
author_facet Lu, Jianglin
Wang, Hailing
Yang, Kuo
Zhang, Yitian
Jenni, Simon
Fu, Yun
contents Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra's Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra's Net. We formalize this hypothesis using the V-enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure-preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross-model and cross-modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training-free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04496
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Indra Representation Hypothesis for Multimodal Alignment
Lu, Jianglin
Wang, Hailing
Yang, Kuo
Zhang, Yitian
Jenni, Simon
Fu, Yun
Computer Vision and Pattern Recognition
Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra's Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra's Net. We formalize this hypothesis using the V-enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure-preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross-model and cross-modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training-free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra.
title The Indra Representation Hypothesis for Multimodal Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04496