Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wenrui, Wang, Penghong, Wang, Xingtao, Zuo, Wangmeng, Fan, Xiaopeng, Tian, Yonghong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
by: Li, Wenrui, et al.
Published: (2026)
by: Li, Wenrui, et al.
Published: (2026)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
PAGCNet: A Pose-Aware and Geometry Constrained Framework for Panoramic Depth Estimation
by: Ning, Kanglin, et al.
Published: (2025)
by: Ning, Kanglin, et al.
Published: (2025)
RoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware Roaming
by: Chu, Jisheng, et al.
Published: (2026)
by: Chu, Jisheng, et al.
Published: (2026)
RouteWinFormer: A Route-Window Transformer for Middle-range Attention in Image Restoration
by: Li, Qifan, et al.
Published: (2025)
by: Li, Qifan, et al.
Published: (2025)
Towards Accurate Single Panoramic 3D Detection: A Semantic Gaussian Centric Approach
by: Ning, Kanglin, et al.
Published: (2026)
by: Ning, Kanglin, et al.
Published: (2026)
Digging into Intrinsic Contextual Information for High-fidelity 3D Point Cloud Completion
by: Chu, Jisheng, et al.
Published: (2024)
by: Chu, Jisheng, et al.
Published: (2024)
Hyperbolic-constraint Point Cloud Reconstruction from Single RGB-D Images
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025)
by: Wang, Zhitao, et al.
Published: (2025)
Bidirectional Feature-aligned Motion Transformation for Efficient Dynamic Point Cloud Compression
by: Deng, Xuan, et al.
Published: (2025)
by: Deng, Xuan, et al.
Published: (2025)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
by: Yu, RunLin, et al.
Published: (2024)
by: Yu, RunLin, et al.
Published: (2024)
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
by: Ning, Kanglin, et al.
Published: (2026)
by: Ning, Kanglin, et al.
Published: (2026)
Multi-modal Crowd Counting via a Broker Modality
by: Meng, Haoliang, et al.
Published: (2024)
by: Meng, Haoliang, et al.
Published: (2024)
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Learning Spatially Decoupled Color Representations for Facial Image Colorization
by: Zhu, Hangyan, et al.
Published: (2024)
by: Zhu, Hangyan, et al.
Published: (2024)
MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
HDI-Former: Hybrid Dynamic Interaction ANN-SNN Transformer for Object Detection Using Frames and Events
by: Li, Dianze, et al.
Published: (2024)
by: Li, Dianze, et al.
Published: (2024)
LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer
by: Chen, Changgu, et al.
Published: (2025)
by: Chen, Changgu, et al.
Published: (2025)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
by: Kurzendörfer, David, et al.
Published: (2024)
by: Kurzendörfer, David, et al.
Published: (2024)
Distributed Zero-Shot Learning for Visual Recognition
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis
by: Wan, Yecong, et al.
Published: (2026)
by: Wan, Yecong, et al.
Published: (2026)
ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
by: Jiang, Kui, et al.
Published: (2025)
by: Jiang, Kui, et al.
Published: (2025)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024)
by: Hou, Wenjin, et al.
Published: (2024)
QKFormer: Hierarchical Spiking Transformer using Q-K Attention
by: Zhou, Chenlin, et al.
Published: (2024)
by: Zhou, Chenlin, et al.
Published: (2024)
PVINet: Point-Voxel Interlaced Network for Point Cloud Compression
by: Deng, Xuan, et al.
Published: (2025)
by: Deng, Xuan, et al.
Published: (2025)
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
by: Li, Yuanze, et al.
Published: (2024)
by: Li, Yuanze, et al.
Published: (2024)
Learning Visual Proxy for Compositional Zero-Shot Learning
by: Zhang, Shiyu, et al.
Published: (2025)
by: Zhang, Shiyu, et al.
Published: (2025)
Responsible Visual Editing
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
by: Cai, Wenrui, et al.
Published: (2026)
by: Cai, Wenrui, et al.
Published: (2026)
Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric
by: Wan, Zhaolin, et al.
Published: (2025)
by: Wan, Zhaolin, et al.
Published: (2025)
PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting
by: Li, Hantang, et al.
Published: (2026)
by: Li, Hantang, et al.
Published: (2026)
VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
by: Hou, Yanning, et al.
Published: (2026)
by: Hou, Yanning, et al.
Published: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Similar Items
-
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024) -
Language-Guided Graph Representation Learning for Video Summarization
by: Li, Wenrui, et al.
Published: (2025) -
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024) -
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
by: Li, Wenrui, et al.
Published: (2026) -
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)