Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenrui, Wang, Penghong, Wang, Xingtao, Zuo, Wangmeng, Fan, Xiaopeng, Tian, Yonghong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913859572334592
author Li, Wenrui
Wang, Penghong
Wang, Xingtao
Zuo, Wangmeng
Fan, Xiaopeng
Tian, Yonghong
author_facet Li, Wenrui
Wang, Penghong
Wang, Xingtao
Zuo, Wangmeng
Fan, Xiaopeng
Tian, Yonghong
contents Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scene biases and inadequate motion detail. This paper proposes a novel dual-stream Multi-Timescale Motion-Decoupled Spiking Transformer (MDST++), which decouples contextual semantic information and sparse dynamic motion information. The recurrent joint learning unit is proposed to extract contextual semantic information and capture joint knowledge across various modalities to understand the environment of actions. By converting RGB images to events, our method captures motion information more accurately and mitigates background scene biases. Moreover, we introduce a discrepancy analysis block to model audio motion information. To enhance the robustness of SNNs in extracting temporal and motion cues, we dynamically adjust the threshold of Leaky Integrate-and-Fire neurons based on global motion and contextual semantic information. Our experiments validate the effectiveness of MDST++, demonstrating their consistent superiority over state-of-the-art methods on mainstream benchmarks. Additionally, incorporating motion and multi-timescale information significantly improves HM and ZSL accuracy by 26.2\% and 39.9\%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Li, Wenrui
Wang, Penghong
Wang, Xingtao
Zuo, Wangmeng
Fan, Xiaopeng
Tian, Yonghong
Computer Vision and Pattern Recognition
Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scene biases and inadequate motion detail. This paper proposes a novel dual-stream Multi-Timescale Motion-Decoupled Spiking Transformer (MDST++), which decouples contextual semantic information and sparse dynamic motion information. The recurrent joint learning unit is proposed to extract contextual semantic information and capture joint knowledge across various modalities to understand the environment of actions. By converting RGB images to events, our method captures motion information more accurately and mitigates background scene biases. Moreover, we introduce a discrepancy analysis block to model audio motion information. To enhance the robustness of SNNs in extracting temporal and motion cues, we dynamically adjust the threshold of Leaky Integrate-and-Fire neurons based on global motion and contextual semantic information. Our experiments validate the effectiveness of MDST++, demonstrating their consistent superiority over state-of-the-art methods on mainstream benchmarks. Additionally, incorporating motion and multi-timescale information significantly improves HM and ZSL accuracy by 26.2\% and 39.9\%.
title Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.19938