FusionSAM: Visual Multi-Modal Learning with Segment Anything

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Daixun, Xie, Weiying, Cao, Mingxiang, Wang, Yunke, Zhang, Yusi, Fang, Leyuan, Li, Yunsong, Xu, Chang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912445367320576
author Li, Daixun
Xie, Weiying
Cao, Mingxiang
Wang, Yunke
Zhang, Yusi
Fang, Leyuan
Li, Yunsong
Xu, Chang
author_facet Li, Daixun
Xie, Weiying
Cao, Mingxiang
Wang, Yunke
Zhang, Yusi
Fang, Leyuan
Li, Yunsong
Xu, Chang
contents Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1$\%$ over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13980
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FusionSAM: Visual Multi-Modal Learning with Segment Anything
Li, Daixun
Xie, Weiying
Cao, Mingxiang
Wang, Yunke
Zhang, Yusi
Fang, Leyuan
Li, Yunsong
Xu, Chang
Computer Vision and Pattern Recognition
Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1$\%$ over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes.
title FusionSAM: Visual Multi-Modal Learning with Segment Anything
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.13980