SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghalkha, Abdulmomen, Tian, Zhuojun, Issaid, Chaouki Ben, Bennis, Mehdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914109164879872
author Ghalkha, Abdulmomen
Tian, Zhuojun
Issaid, Chaouki Ben
Bennis, Mehdi
author_facet Ghalkha, Abdulmomen
Tian, Zhuojun
Issaid, Chaouki Ben
Bennis, Mehdi
contents Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment
Ghalkha, Abdulmomen
Tian, Zhuojun
Issaid, Chaouki Ben
Bennis, Mehdi
Machine Learning
Conventional multimodal alignment methods assume mutual redundancy across all modalities, an assumption that fails in real-world distributed scenarios. We propose SheafAlign, a sheaf-theoretic framework for decentralized multimodal alignment that replaces single-space alignment with multiple comparison spaces. This approach models pairwise modality relations through sheaf structures and leverages decentralized contrastive learning-based objectives for training. SheafAlign overcomes the limitations of prior methods by not requiring mutual redundancy among all modalities, preserving both shared and unique information. Experiments on multimodal sensing datasets show superior zero-shot generalization, cross-modal alignment, and robustness to missing modalities, with 50\% lower communication cost than state-of-the-art baselines.
title SheafAlign: A Sheaf-theoretic Framework for Decentralized Multimodal Alignment
topic Machine Learning
url https://arxiv.org/abs/2510.20540