Reconstruction as a Bridge for Event-Based Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917141887844352 |
|---|---|
| author | Lou, Hanyue Zhou, Jiayi Zhang, Yang Li, Boyu Wang, Yi Ye, Guangnan Shi, Boxin |
| author_facet | Lou, Hanyue Zhou, Jiayi Zhang, Yang Li, Boyu Wang, Yi Ye, Guangnan Shi, Boxin |
| contents | Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_11510 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Reconstruction as a Bridge for Event-Based Visual Question Answering Lou, Hanyue Zhou, Jiayi Zhang, Yang Li, Boyu Wang, Yi Ye, Guangnan Shi, Boxin Computer Vision and Pattern Recognition Integrating event cameras with Multimodal Large Language Models (MLLMs) promises general scene understanding in challenging visual conditions, yet requires navigating a trade-off between preserving the unique advantages of event data and ensuring compatibility with frame-based models. We address this challenge by using reconstruction as a bridge, proposing a straightforward Frame-based Reconstruction and Tokenization (FRT) method and designing an efficient Adaptive Reconstruction and Tokenization (ART) method that leverages event sparsity. For robust evaluation, we introduce EvQA, the first objective, real-world benchmark for event-based MLLMs, comprising 1,000 event-Q&A pairs from 22 public datasets. Our experiments demonstrate that our methods achieve state-of-the-art performance on EvQA, highlighting the significant potential of MLLMs in event-based vision. |
| title | Reconstruction as a Bridge for Event-Based Visual Question Answering |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.11510 |