CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Broedermann, Tim, Sakaridis, Christos, Fu, Yuqian, Van Gool, Luc
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915122038964224
author Broedermann, Tim
Sakaridis, Christos
Fu, Yuqian
Van Gool, Luc
author_facet Broedermann, Tim
Sakaridis, Christos
Fu, Yuqian
Van Gool, Luc
contents Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing sensor fusion methods often treat sensors uniformly across all conditions, leading to suboptimal performance. By contrast, we propose a novel, condition-aware multimodal fusion approach for robust semantic perception of driving scenes. Our method, CAFuser, uses an RGB camera input to classify environmental conditions and generate a Condition Token that guides the fusion of multiple sensor modalities. We further newly introduce modality-specific feature adapters to align diverse sensor inputs into a shared latent space, enabling efficient integration with a single and shared pre-trained backbone. By dynamically adapting sensor fusion based on the actual condition, our model significantly improves robustness and accuracy, especially in adverse-condition scenarios. CAFuser ranks first on the public MUSES benchmarks, achieving 59.7 PQ for multimodal panoptic and 78.2 mIoU for semantic segmentation, and also sets the new state of the art on DeLiVER. The source code is publicly available at: https://github.com/timbroed/CAFuser.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10791
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
Broedermann, Tim
Sakaridis, Christos
Fu, Yuqian
Van Gool, Luc
Computer Vision and Pattern Recognition
Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing sensor fusion methods often treat sensors uniformly across all conditions, leading to suboptimal performance. By contrast, we propose a novel, condition-aware multimodal fusion approach for robust semantic perception of driving scenes. Our method, CAFuser, uses an RGB camera input to classify environmental conditions and generate a Condition Token that guides the fusion of multiple sensor modalities. We further newly introduce modality-specific feature adapters to align diverse sensor inputs into a shared latent space, enabling efficient integration with a single and shared pre-trained backbone. By dynamically adapting sensor fusion based on the actual condition, our model significantly improves robustness and accuracy, especially in adverse-condition scenarios. CAFuser ranks first on the public MUSES benchmarks, achieving 59.7 PQ for multimodal panoptic and 78.2 mIoU for semantic segmentation, and also sets the new state of the art on DeLiVER. The source code is publicly available at: https://github.com/timbroed/CAFuser.
title CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.10791