VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911149220429824 |
|---|---|
| author | Le, Chenqian Zhao, Yilin Emami, Nikasadat Yadav, Kushagra Liu, Xujin "Chris" Chen, Xupeng Wang, Yao |
| author_facet | Le, Chenqian Zhao, Yilin Emami, Nikasadat Yadav, Kushagra Liu, Xujin "Chris" Chen, Xupeng Wang, Yao |
| contents | Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scalability and practical deployment. We introduce \textbf{VoxelFormer}, a lightweight transformer architecture that enables multi-subject training for visual decoding from fMRI. VoxelFormer integrates a Token Merging Transformer (ToMer) for efficient voxel compression and a query-driven Q-Former that produces fixed-size neural representations aligned with the CLIP image embedding space. Evaluated on the 7T Natural Scenes Dataset, VoxelFormer achieves competitive retrieval performance on subjects included during training with significantly fewer parameters than existing methods. These results highlight token merging and query-based transformers as promising strategies for parameter-efficient neural decoding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_09015 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI Le, Chenqian Zhao, Yilin Emami, Nikasadat Yadav, Kushagra Liu, Xujin "Chris" Chen, Xupeng Wang, Yao Computer Vision and Pattern Recognition Recent advances in fMRI-based visual decoding have enabled compelling reconstructions of perceived images. However, most approaches rely on subject-specific training, limiting scalability and practical deployment. We introduce \textbf{VoxelFormer}, a lightweight transformer architecture that enables multi-subject training for visual decoding from fMRI. VoxelFormer integrates a Token Merging Transformer (ToMer) for efficient voxel compression and a query-driven Q-Former that produces fixed-size neural representations aligned with the CLIP image embedding space. Evaluated on the 7T Natural Scenes Dataset, VoxelFormer achieves competitive retrieval performance on subjects included during training with significantly fewer parameters than existing methods. These results highlight token merging and query-based transformers as promising strategies for parameter-efficient neural decoding. |
| title | VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.09015 |