Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Shiyi, Liang, Dong, Zheng, Hairong, Zhou, Yihang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2510.03122
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914305291583488
author Zhang, Shiyi
Liang, Dong
Zheng, Hairong
Zhou, Yihang
author_facet Zhang, Shiyi
Liang, Dong
Zheng, Hairong
Zhou, Yihang
contents The reconstruction of visual information from brain activity fosters interdisciplinary integration between neuroscience and computer vision. However, existing methods still face challenges in accurately recovering highly complex visual stimuli. This difficulty stems from the characteristics of natural scenes: low-level features exhibit heterogeneity, while high-level features show semantic entanglement due to contextual overlaps. Inspired by the hierarchical representation theory of the visual cortex, we propose the HAVIR model, which separates the visual cortex into two hierarchical regions and extracts distinct features from each. Specifically, the Structural Generator extracts structural information from spatial processing voxels and converts it into latent diffusion priors, while the Semantic Extractor converts semantic processing voxels into CLIP embeddings. These components are integrated via the Versatile Diffusion model to synthesize the final image. Experimental results demonstrate that HAVIR enhances both the structural and semantic quality of reconstructions, even in complex scenes, and outperforms existing models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
Zhang, Shiyi
Liang, Dong
Zheng, Hairong
Zhou, Yihang
Computer Vision and Pattern Recognition
Artificial Intelligence
The reconstruction of visual information from brain activity fosters interdisciplinary integration between neuroscience and computer vision. However, existing methods still face challenges in accurately recovering highly complex visual stimuli. This difficulty stems from the characteristics of natural scenes: low-level features exhibit heterogeneity, while high-level features show semantic entanglement due to contextual overlaps. Inspired by the hierarchical representation theory of the visual cortex, we propose the HAVIR model, which separates the visual cortex into two hierarchical regions and extracts distinct features from each. Specifically, the Structural Generator extracts structural information from spatial processing voxels and converts it into latent diffusion priors, while the Semantic Extractor converts semantic processing voxels into CLIP embeddings. These components are integrated via the Versatile Diffusion model to synthesize the final image. Experimental results demonstrate that HAVIR enhances both the structural and semantic quality of reconstructions, even in complex scenes, and outperforms existing models.
title HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.03122