V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916752466640896 |
|---|---|
| author | Lou, Hanyue Liang, Jinxiu Teng, Minggui Wang, Yi Shi, Boxin |
| author_facet | Lou, Hanyue Liang, Jinxiu Teng, Minggui Wang, Yi Shi, Boxin |
| contents | Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16797 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation Lou, Hanyue Liang, Jinxiu Teng, Minggui Wang, Yi Shi, Boxin Computer Vision and Pattern Recognition Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the scarcity of real data prevent event-based training datasets from scaling up, limiting the development and generalization capabilities of event vision models. To address this challenge, we introduce Video-to-Voxel (V2V), an approach that directly converts conventional video frames into event-based voxel grid representations, bypassing the storage-intensive event stream generation entirely. V2V enables a 150 times reduction in storage requirements while supporting on-the-fly parameter randomization for enhanced model robustness. Leveraging this efficiency, we train several video reconstruction and optical flow estimation model architectures on 10,000 diverse videos totaling 52 hours--an order of magnitude larger than existing event datasets, yielding substantial improvements. |
| title | V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.16797 |