Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: You, Junwon, Jang, Mihyun, Mo, Sangwoo, Jung, Jae-Hun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909000595931136
author You, Junwon
Jang, Mihyun
Mo, Sangwoo
Jung, Jae-Hun
author_facet You, Junwon
Jang, Mihyun
Mo, Sangwoo
Jung, Jae-Hun
contents Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs together with abundant unlabeled images, existing methods remain fundamentally pairwise and fail to model the global structure of multimodal representation manifolds. Existing topology-based alignment methods rely on persistence diagram matching, which neither guarantees geometric alignment nor utilizes the image-text pairing information central to vision-language learning. We propose Topology-Aware Multimodal Representation Alignment (ToMA), a framework that uses persistent homology to identify topologically salient edges and aligns them across modalities through available cross-modal correspondences. ToMA leverages both H_0-death edges and lightweight H_1-birth edges, allowing it to capture both connectivity and cycle structure without constructing 2-simplices. Experiments show that ToMA yields stable gains, with clear improvements on remote sensing and modest but consistent benefits on fashion retrieval. Additional analysis shows that ToMA is more stable than alternative topology-based objectives and that lightweight H_1-birth edges provide useful higher-order structural signals.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26370
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning
You, Junwon
Jang, Mihyun
Mo, Sangwoo
Jung, Jae-Hun
Computer Vision and Pattern Recognition
Machine Learning
Algebraic Topology
Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs together with abundant unlabeled images, existing methods remain fundamentally pairwise and fail to model the global structure of multimodal representation manifolds. Existing topology-based alignment methods rely on persistence diagram matching, which neither guarantees geometric alignment nor utilizes the image-text pairing information central to vision-language learning. We propose Topology-Aware Multimodal Representation Alignment (ToMA), a framework that uses persistent homology to identify topologically salient edges and aligns them across modalities through available cross-modal correspondences. ToMA leverages both H_0-death edges and lightweight H_1-birth edges, allowing it to capture both connectivity and cycle structure without constructing 2-simplices. Experiments show that ToMA yields stable gains, with clear improvements on remote sensing and modest but consistent benefits on fashion retrieval. Additional analysis shows that ToMA is more stable than alternative topology-based objectives and that lightweight H_1-birth edges provide useful higher-order structural signals.
title Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning
topic Computer Vision and Pattern Recognition
Machine Learning
Algebraic Topology
url https://arxiv.org/abs/2604.26370