Reconstruction as a Bridge for Event-Based Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lou, Hanyue, Zhou, Jiayi, Zhang, Yang, Li, Boyu, Wang, Yi, Ye, Guangnan, Shi, Boxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917141887844352
author Lou, Hanyue
Zhou, Jiayi
Zhang, Yang
Li, Boyu
Wang, Yi
Ye, Guangnan
Shi, Boxin
author_facet Lou, Hanyue
Zhou, Jiayi
Zhang, Yang
Li, Boyu
Wang, Yi
Ye, Guangnan
Shi, Boxin
contents Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11510
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reconstruction as a Bridge for Event-Based Visual Question Answering
Lou, Hanyue
Zhou, Jiayi
Zhang, Yang
Li, Boyu
Wang, Yi
Ye, Guangnan
Shi, Boxin
Computer Vision and Pattern Recognition
Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision.
title Reconstruction as a Bridge for Event-Based Visual Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11510