SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Gaole, Dong, Menghang, Zhang, Rongyu, An, Ruichuan, Zhang, Shanghang, Huang, Tiejun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909817257328640
author Dai, Gaole
Dong, Menghang
Zhang, Rongyu
An, Ruichuan
Zhang, Shanghang
Huang, Tiejun
author_facet Dai, Gaole
Dong, Menghang
Zhang, Rongyu
An, Ruichuan
Zhang, Shanghang
Huang, Tiejun
contents The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the functions of cone and rod cells: cones are primarily responsible for color perception, while rods are specialized in detecting motion, particularly variations in light intensity. These two distinct modalities of visual information are integrated and processed within the visual cortex, thereby enhancing the robustness of the human visual system. Inspired by this biological mechanism, modern hardware systems have evolved to include not only color-sensitive RGB cameras but also motion-sensitive Dynamic Visual Systems, such as spike cameras. Building upon these advancements, this study seeks to emulate the human visual system by integrating decomposed multi-modal visual inputs with modern latent-space generative frameworks. We named it SpikeGen. We evaluate its performance across various spike-RGB tasks, including conditional image and video deblurring, dense frame reconstruction from spike streams, and high-speed scene novel-view synthesis. Supported by extensive experiments, we demonstrate that leveraging the latent space manipulation capabilities of generative models enables an effective synergistic enhancement of different visual modalities, addressing spatial sparsity in spike inputs and temporal sparsity in RGB inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
Dai, Gaole
Dong, Menghang
Zhang, Rongyu
An, Ruichuan
Zhang, Shanghang
Huang, Tiejun
Computer Vision and Pattern Recognition
The process through which humans perceive and learn visual representations in dynamic environments is highly complex. From a structural perspective, the human eye decouples the functions of cone and rod cells: cones are primarily responsible for color perception, while rods are specialized in detecting motion, particularly variations in light intensity. These two distinct modalities of visual information are integrated and processed within the visual cortex, thereby enhancing the robustness of the human visual system. Inspired by this biological mechanism, modern hardware systems have evolved to include not only color-sensitive RGB cameras but also motion-sensitive Dynamic Visual Systems, such as spike cameras. Building upon these advancements, this study seeks to emulate the human visual system by integrating decomposed multi-modal visual inputs with modern latent-space generative frameworks. We named it SpikeGen. We evaluate its performance across various spike-RGB tasks, including conditional image and video deblurring, dense frame reconstruction from spike streams, and high-speed scene novel-view synthesis. Supported by extensive experiments, we demonstrate that leveraging the latent space manipulation capabilities of generative models enables an effective synergistic enhancement of different visual modalities, addressing spatial sparsity in spike inputs and temporal sparsity in RGB inputs.
title SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18049