Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Ruiping, Zhang, Jiaming, Peng, Kunyu, Chen, Yufan, Cao, Ke, Zheng, Junwei, Sarfraz, M. Saquib, Yang, Kailun, Stiefelhagen, Rainer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911835314192384
author Liu, Ruiping
Zhang, Jiaming
Peng, Kunyu
Chen, Yufan
Cao, Ke
Zheng, Junwei
Sarfraz, M. Saquib
Yang, Kailun
Stiefelhagen, Rainer
author_facet Liu, Ruiping
Zhang, Jiaming
Peng, Kunyu
Chen, Yufan
Cao, Ke
Zheng, Junwei
Sarfraz, M. Saquib
Yang, Kailun
Stiefelhagen, Rainer
contents Integrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework. However, the modality incompleteness in multi-modal segmentation remains under-explored. In this work, we establish a task called Modality-Incomplete Scene Segmentation (MISS), which encompasses both system-level modality absence and sensor-level modality errors. To avoid the predominant modality reliance in multi-modal fusion, we introduce a Missing-aware Modal Switch (MMS) strategy to proactively manage missing modalities during training. Utilizing bit-level batch-wise sampling enhances the model's performance in both complete and incomplete testing scenarios. Furthermore, we introduce the Fourier Prompt Tuning (FPT) method to incorporate representative spectral information into a limited number of learnable prompts that maintain robustness against all MISS scenarios. Akin to fine-tuning effects but with fewer tunable parameters (1.1%). Extensive experiments prove the efficacy of our proposed approach, showcasing an improvement of 5.84% mIoU over the prior state-of-the-art parameter-efficient methods in modality missing. The source code is publicly available at https://github.com/RuipingL/MISS.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
Liu, Ruiping
Zhang, Jiaming
Peng, Kunyu
Chen, Yufan
Cao, Ke
Zheng, Junwei
Sarfraz, M. Saquib
Yang, Kailun
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
Integrating information from multiple modalities enhances the robustness of scene perception systems in autonomous vehicles, providing a more comprehensive and reliable sensory framework. However, the modality incompleteness in multi-modal segmentation remains under-explored. In this work, we establish a task called Modality-Incomplete Scene Segmentation (MISS), which encompasses both system-level modality absence and sensor-level modality errors. To avoid the predominant modality reliance in multi-modal fusion, we introduce a Missing-aware Modal Switch (MMS) strategy to proactively manage missing modalities during training. Utilizing bit-level batch-wise sampling enhances the model's performance in both complete and incomplete testing scenarios. Furthermore, we introduce the Fourier Prompt Tuning (FPT) method to incorporate representative spectral information into a limited number of learnable prompts that maintain robustness against all MISS scenarios. Akin to fine-tuning effects but with fewer tunable parameters (1.1%). Extensive experiments prove the efficacy of our proposed approach, showcasing an improvement of 5.84% mIoU over the prior state-of-the-art parameter-efficient methods in modality missing. The source code is publicly available at https://github.com/RuipingL/MISS.
title Fourier Prompt Tuning for Modality-Incomplete Scene Segmentation
topic Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
url https://arxiv.org/abs/2401.16923