NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Zixuan, Bao, Guangyin, Zhang, Qi, Wan, Zhongwei, Miao, Duoqian, Wang, Shoujin, Zhu, Lei, Wang, Changwei, Xu, Rongtao, Hu, Liang, Liu, Ke, Zhang, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912155605925888
author Gong, Zixuan
Bao, Guangyin
Zhang, Qi
Wan, Zhongwei
Miao, Duoqian
Wang, Shoujin
Zhu, Lei
Wang, Changwei
Xu, Rongtao
Hu, Liang
Liu, Ke
Zhang, Yu
author_facet Gong, Zixuan
Bao, Guangyin
Zhang, Qi
Wan, Zhongwei
Miao, Duoqian
Wang, Shoujin
Zhu, Lei
Wang, Changwei
Xu, Rongtao
Hu, Liang
Liu, Ke
Zhang, Yu
contents Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19452
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
Gong, Zixuan
Bao, Guangyin
Zhang, Qi
Wan, Zhongwei
Miao, Duoqian
Wang, Shoujin
Zhu, Lei
Wang, Changwei
Xu, Rongtao
Hu, Liang
Liu, Ke
Zhang, Yu
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
title NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.19452