SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wenrui, Hong, Xiaopeng, Xiong, Ruiqin, Fan, Xiaopeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025)
by: Wang, Zhitao, et al.
Published: (2025)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and Trajectory Information
by: Li, Jie, et al.
Published: (2023)
by: Li, Jie, et al.
Published: (2023)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
by: Moradi, Morteza, et al.
Published: (2024)
by: Moradi, Morteza, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
by: Wang, Xinghan, et al.
Published: (2024)
by: Wang, Xinghan, et al.
Published: (2024)
Using Saliency and Cropping to Improve Video Memorability
by: Mudgal, Vaibhav, et al.
Published: (2023)
by: Mudgal, Vaibhav, et al.
Published: (2023)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
by: Zhu, Yixin, et al.
Published: (2026)
by: Zhu, Yixin, et al.
Published: (2026)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
Robust Mesh Saliency Ground Truth Acquisition in VR via View Cone Sampling and Manifold Diffusion
by: Zheng, Guoquan, et al.
Published: (2026)
by: Zheng, Guoquan, et al.
Published: (2026)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
by: Fan, Xinqi, et al.
Published: (2025)
by: Fan, Xinqi, et al.
Published: (2025)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
by: Zhu, Haodong, et al.
Published: (2025)
by: Zhu, Haodong, et al.
Published: (2025)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
by: Kandhare, Mahesh, et al.
Published: (2024)
by: Kandhare, Mahesh, et al.
Published: (2024)
Towards Universal Modal Tracking with Online Dense Temporal Token Learning
by: Zheng, Yaozong, et al.
Published: (2025)
by: Zheng, Yaozong, et al.
Published: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
Boosting Temporal Sentence Grounding via Causal Inference
by: Tang, Kefan, et al.
Published: (2025)
by: Tang, Kefan, et al.
Published: (2025)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
by: Song, Shezheng, et al.
Published: (2023)
by: Song, Shezheng, et al.
Published: (2023)
Anisotropic Modality Align
by: Yu, Xiaomin, et al.
Published: (2026)
by: Yu, Xiaomin, et al.
Published: (2026)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
by: Li, Chengzhi, et al.
Published: (2025)
by: Li, Chengzhi, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
by: Meng, Jiahao, et al.
Published: (2026)
by: Meng, Jiahao, et al.
Published: (2026)
Similar Items
-
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024) -
Spiking Variational Graph Representation Inference for Video Summarization
by: Li, Wenrui, et al.
Published: (2025) -
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
by: Wang, Zhitao, et al.
Published: (2025) -
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024) -
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)