CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zelin, Li, Kedi, Liang, Huiqi, Zhang, Tao, Xu, Chuanzhi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908934839730176
author Zhang, Zelin
Li, Kedi
Liang, Huiqi
Zhang, Tao
Xu, Chuanzhi
author_facet Zhang, Zelin
Li, Kedi
Liang, Huiqi
Zhang, Tao
Xu, Chuanzhi
contents Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use modality-specific adaptations or rely on loosely coupled interactions, thereby limiting flexibility and resulting in less effective cross-modal coordination. Moreover, these methods often struggle to balance efficient information exchange with preserving the unique characteristics of each modality across different modality combinations. To address these challenges, we propose CrossWeaver, a simple yet effective multimodal fusion framework for arbitrary-modality semantic segmentation. Its core is a Modality Interaction Block (MIB), which enables selective and reliability-aware cross-modal interaction within the encoder, while a lightweight Seam-Aligned Fusion (SAF) module further aggregates the enhanced features. Extensive experiments on multiple multimodal semantic segmentation benchmarks demonstrate that our framework achieves state-of-the-art performance with minimal additional parameters and strong generalization to unseen modality combinations.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02948
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
Zhang, Zelin
Li, Kedi
Liang, Huiqi
Zhang, Tao
Xu, Chuanzhi
Computer Vision and Pattern Recognition
Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use modality-specific adaptations or rely on loosely coupled interactions, thereby limiting flexibility and resulting in less effective cross-modal coordination. Moreover, these methods often struggle to balance efficient information exchange with preserving the unique characteristics of each modality across different modality combinations. To address these challenges, we propose CrossWeaver, a simple yet effective multimodal fusion framework for arbitrary-modality semantic segmentation. Its core is a Modality Interaction Block (MIB), which enables selective and reliability-aware cross-modal interaction within the encoder, while a lightweight Seam-Aligned Fusion (SAF) module further aggregates the enhanced features. Extensive experiments on multiple multimodal semantic segmentation benchmarks demonstrate that our framework achieves state-of-the-art performance with minimal additional parameters and strong generalization to unseen modality combinations.
title CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.02948