InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakhaia, Mustafa, Sithua, Kaung, Okea, Min Khant Soe, Wielgosza, Maciej
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909016473468928
author Sakhaia, Mustafa
Sithua, Kaung
Okea, Min Khant Soe
Wielgosza, Maciej
author_facet Sakhaia, Mustafa
Sithua, Kaung
Okea, Min Khant Soe
Wielgosza, Maciej
contents Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often struggle in high-dynamic- range scenes or high-speed scenarios due to motion blur and latency. Dynamic Vision Sensors (DVS), or event cameras, offer a paradigm shift by capturing asynchronous brightness changes with microsecond temporal resolution and high dynamic range. In this paper, we propose an extended architecture of the state-of-the-art InterFuser model, integrating DVS as an additional modality to enhance perception reliability. We introduce a novel token-based fusion strategy that incorporates accumulated event frames into the transformer-based backbone of InterFuser. Our method leverages the complementary nature of RGB, LiDAR, and DVS data. We evaluate our approach on the Car Learning to Act (CARLA) Leaderboard benchmarks, demonstrating that the inclusion of DVS improves the robustness of the driving agent, achieving a competitive Driving Score of 77.2 and a superior Route Completion of 100%. The results indicate that event-based vision is a promising direction for improving safety and performance in adverse lighting and dynamic conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04355
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making
Sakhaia, Mustafa
Sithua, Kaung
Okea, Min Khant Soe
Wielgosza, Maciej
Computer Vision and Pattern Recognition
Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often struggle in high-dynamic- range scenes or high-speed scenarios due to motion blur and latency. Dynamic Vision Sensors (DVS), or event cameras, offer a paradigm shift by capturing asynchronous brightness changes with microsecond temporal resolution and high dynamic range. In this paper, we propose an extended architecture of the state-of-the-art InterFuser model, integrating DVS as an additional modality to enhance perception reliability. We introduce a novel token-based fusion strategy that incorporates accumulated event frames into the transformer-based backbone of InterFuser. Our method leverages the complementary nature of RGB, LiDAR, and DVS data. We evaluate our approach on the Car Learning to Act (CARLA) Leaderboard benchmarks, demonstrating that the inclusion of DVS improves the robustness of the driving agent, achieving a competitive Driving Score of 77.2 and a superior Route Completion of 100%. The results indicate that event-based vision is a promising direction for improving safety and performance in adverse lighting and dynamic conditions.
title InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.04355