EvDiff: High Quality Video with an Event Camera

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Weilun, Sun, Lei, Gao, Ruixi, Jiang, Qi, Ma, Yuqin, Wang, Kaiwei, Yang, Ming-Hsuan, Van Gool, Luc, Paudel, Danda Pani
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914166636281856
author Li, Weilun
Sun, Lei
Gao, Ruixi
Jiang, Qi
Ma, Yuqin
Wang, Kaiwei
Yang, Ming-Hsuan
Van Gool, Luc
Paudel, Danda Pani
author_facet Li, Weilun
Sun, Lei
Gao, Ruixi
Jiang, Qi
Ma, Yuqin
Wang, Kaiwei
Yang, Ming-Hsuan
Van Gool, Luc
Paudel, Danda Pani
contents As neuromorphic sensors, event cameras asynchronously record changes in brightness as streams of sparse events with the advantages of high temporal resolution and high dynamic range. Reconstructing intensity images from events is a highly ill-posed task due to the inherent ambiguity of absolute brightness. Early methods generally follow an end-to-end regression paradigm, directly mapping events to intensity frames in a deterministic manner. While effective to some extent, these approaches often yield perceptually inferior results and struggle to scale up in model capacity and training data. In this work, we propose EvDiff, an event-based diffusion model that follows a surrogate training framework to produce high-quality videos. To reduce the heavy computational cost of high-frame-rate video generation, we design an event-based diffusion model that performs only a single forward diffusion step, equipped with a temporally consistent EvEncoder. Furthermore, our novel Surrogate Training Framework eliminates the dependence on paired event-image datasets, allowing the model to leverage large-scale image datasets for higher capacity. The proposed EvDiff is capable of generating high-quality colorful videos solely from monochromatic event streams. Experiments on real-world datasets demonstrate that our method strikes a sweet spot between fidelity and realism, outperforming existing approaches on both pixel-level and perceptual metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EvDiff: High Quality Video with an Event Camera
Li, Weilun
Sun, Lei
Gao, Ruixi
Jiang, Qi
Ma, Yuqin
Wang, Kaiwei
Yang, Ming-Hsuan
Van Gool, Luc
Paudel, Danda Pani
Computer Vision and Pattern Recognition
As neuromorphic sensors, event cameras asynchronously record changes in brightness as streams of sparse events with the advantages of high temporal resolution and high dynamic range. Reconstructing intensity images from events is a highly ill-posed task due to the inherent ambiguity of absolute brightness. Early methods generally follow an end-to-end regression paradigm, directly mapping events to intensity frames in a deterministic manner. While effective to some extent, these approaches often yield perceptually inferior results and struggle to scale up in model capacity and training data. In this work, we propose EvDiff, an event-based diffusion model that follows a surrogate training framework to produce high-quality videos. To reduce the heavy computational cost of high-frame-rate video generation, we design an event-based diffusion model that performs only a single forward diffusion step, equipped with a temporally consistent EvEncoder. Furthermore, our novel Surrogate Training Framework eliminates the dependence on paired event-image datasets, allowing the model to leverage large-scale image datasets for higher capacity. The proposed EvDiff is capable of generating high-quality colorful videos solely from monochromatic event streams. Experiments on real-world datasets demonstrate that our method strikes a sweet spot between fidelity and realism, outperforming existing approaches on both pixel-level and perceptual metrics.
title EvDiff: High Quality Video with an Event Camera
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17492