SpikeCLIP: A Contrastive Language-Image Pretrained Spiking Neural Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Changze, Li, Tianlong, Liu, Wenhao, Gu, Yufei, Xu, Jianhan, Zhang, Cenyuan, Wu, Muling, Zheng, Xiaoqing, Huang, Xuanjing
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909617633624064
author Lv, Changze
Li, Tianlong
Liu, Wenhao
Gu, Yufei
Xu, Jianhan
Zhang, Cenyuan
Wu, Muling
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Lv, Changze
Li, Tianlong
Liu, Wenhao
Gu, Yufei
Xu, Jianhan
Zhang, Cenyuan
Wu, Muling
Zheng, Xiaoqing
Huang, Xuanjing
contents Spiking Neural Networks (SNNs) have emerged as a promising alternative to conventional Artificial Neural Networks (ANNs), demonstrating comparable performance in both visual and linguistic tasks while offering the advantage of improved energy efficiency. Despite these advancements, the integration of linguistic and visual features into a unified representation through spike trains poses a significant challenge, and the application of SNNs to multimodal scenarios remains largely unexplored. This paper presents SpikeCLIP, a novel framework designed to bridge the modality gap in spike-based computation. Our approach employs a two-step recipe: an ``alignment pre-training'' to align features across modalities, followed by a ``dual-loss fine-tuning'' to refine the model's performance. Extensive experiments reveal that SNNs achieve results on par with ANNs while substantially reducing energy consumption across various datasets commonly used for multimodal model evaluation. Furthermore, SpikeCLIP maintains robust image classification capabilities, even when dealing with classes that fall outside predefined categories. This study marks a significant advancement in the development of energy-efficient and biologically plausible multimodal learning systems. Our code is available at https://github.com/Lvchangze/SpikeCLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06488
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SpikeCLIP: A Contrastive Language-Image Pretrained Spiking Neural Network
Lv, Changze
Li, Tianlong
Liu, Wenhao
Gu, Yufei
Xu, Jianhan
Zhang, Cenyuan
Wu, Muling
Zheng, Xiaoqing
Huang, Xuanjing
Neural and Evolutionary Computing
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Spiking Neural Networks (SNNs) have emerged as a promising alternative to conventional Artificial Neural Networks (ANNs), demonstrating comparable performance in both visual and linguistic tasks while offering the advantage of improved energy efficiency. Despite these advancements, the integration of linguistic and visual features into a unified representation through spike trains poses a significant challenge, and the application of SNNs to multimodal scenarios remains largely unexplored. This paper presents SpikeCLIP, a novel framework designed to bridge the modality gap in spike-based computation. Our approach employs a two-step recipe: an ``alignment pre-training'' to align features across modalities, followed by a ``dual-loss fine-tuning'' to refine the model's performance. Extensive experiments reveal that SNNs achieve results on par with ANNs while substantially reducing energy consumption across various datasets commonly used for multimodal model evaluation. Furthermore, SpikeCLIP maintains robust image classification capabilities, even when dealing with classes that fall outside predefined categories. This study marks a significant advancement in the development of energy-efficient and biologically plausible multimodal learning systems. Our code is available at https://github.com/Lvchangze/SpikeCLIP.
title SpikeCLIP: A Contrastive Language-Image Pretrained Spiking Neural Network
topic Neural and Evolutionary Computing
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2310.06488