VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Taparia, Aditya, Ngu, Noel, Leiva, Mario, Kricheli, Joshua Shay, Corcoran, John, Bastian, Nathaniel D., Simari, Gerardo, Shakarian, Paulo, Senanayake, Ransalu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916742968639488
author Taparia, Aditya
Ngu, Noel
Leiva, Mario
Kricheli, Joshua Shay
Corcoran, John
Bastian, Nathaniel D.
Simari, Gerardo
Shakarian, Paulo
Senanayake, Ransalu
author_facet Taparia, Aditya
Ngu, Noel
Leiva, Mario
Kricheli, Joshua Shay
Corcoran, John
Bastian, Nathaniel D.
Simari, Gerardo
Shakarian, Paulo
Senanayake, Ransalu
contents Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in environmental conditions and sensor inputs. As a result, they struggle to adaptively weight each modality under such variations. To address this challenge, we introduce Vision-Language Conditioned Fusion (VLC Fusion), a novel fusion framework that leverages a Vision-Language Model (VLM) to condition the fusion process on nuanced environmental cues. By capturing high-level environmental context such as as darkness, rain, and camera blurring, the VLM guides the model to dynamically adjust modality weights based on the current scene. We evaluate VLC Fusion on real-world autonomous driving and military target detection datasets that include image, LIDAR, and mid-wave infrared modalities. Our experiments show that VLC Fusion consistently outperforms conventional fusion baselines, achieving improved detection accuracy in both seen and unseen scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection
Taparia, Aditya
Ngu, Noel
Leiva, Mario
Kricheli, Joshua Shay
Corcoran, John
Bastian, Nathaniel D.
Simari, Gerardo
Shakarian, Paulo
Senanayake, Ransalu
Computer Vision and Pattern Recognition
Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in environmental conditions and sensor inputs. As a result, they struggle to adaptively weight each modality under such variations. To address this challenge, we introduce Vision-Language Conditioned Fusion (VLC Fusion), a novel fusion framework that leverages a Vision-Language Model (VLM) to condition the fusion process on nuanced environmental cues. By capturing high-level environmental context such as as darkness, rain, and camera blurring, the VLM guides the model to dynamically adjust modality weights based on the current scene. We evaluate VLC Fusion on real-world autonomous driving and military target detection datasets that include image, LIDAR, and mid-wave infrared modalities. Our experiments show that VLC Fusion consistently outperforms conventional fusion baselines, achieving improved detection accuracy in both seen and unseen scenarios.
title VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12715