V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lou, Hanyue, Liang, Jinxiu, Teng, Minggui, Wang, Yi, Shi, Boxin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916752466640896
author Lou, Hanyue
Liang, Jinxiu
Teng, Minggui
Wang, Yi
Shi, Boxin
author_facet Lou, Hanyue
Liang, Jinxiu
Teng, Minggui
Wang, Yi
Shi, Boxin
contents Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
Lou, Hanyue
Liang, Jinxiu
Teng, Minggui
Wang, Yi
Shi, Boxin
Computer Vision and Pattern Recognition
Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements.
title V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16797