State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jiahuan, Zhu, Kai, Cui, Zhenyu, Liu, Zichen, Zou, Xu, Hua, Gang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token Coordinated Prompt Attention is Needed for Visual Prompting
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
by: Zhou, Jiahuan, et al.
Published: (2025)
by: Zhou, Jiahuan, et al.
Published: (2025)
RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
by: Zhu, Kai, et al.
Published: (2026)
by: Zhu, Kai, et al.
Published: (2026)
GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision Model
by: Ai, Zixiang, et al.
Published: (2025)
by: Ai, Zixiang, et al.
Published: (2025)
Componential Prompt-Knowledge Alignment for Domain Incremental Learning
by: Xu, Kunlun, et al.
Published: (2025)
by: Xu, Kunlun, et al.
Published: (2025)
Selective Visual Prompting in Vision Mamba
by: Yao, Yifeng, et al.
Published: (2024)
by: Yao, Yifeng, et al.
Published: (2024)
Vision Graph Prompting via Semantic Low-Rank Decomposition
by: Ai, Zixiang, et al.
Published: (2025)
by: Ai, Zixiang, et al.
Published: (2025)
SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis
by: Ai, Zixiang, et al.
Published: (2025)
by: Ai, Zixiang, et al.
Published: (2025)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Bi-C2R: Bidirectional Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification
by: Cui, Zhenyu, et al.
Published: (2025)
by: Cui, Zhenyu, et al.
Published: (2025)
CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
by: Cui, Zhenyu, et al.
Published: (2025)
by: Cui, Zhenyu, et al.
Published: (2025)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
by: Yang, Min, et al.
Published: (2024)
by: Yang, Min, et al.
Published: (2024)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
AI-Generated Video Detection via Spatio-Temporal Anomaly Learning
by: Bai, Jianfa, et al.
Published: (2024)
by: Bai, Jianfa, et al.
Published: (2024)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
by: Lee, Yeonkyung, et al.
Published: (2026)
by: Lee, Yeonkyung, et al.
Published: (2026)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
by: Ge, Chengjie, et al.
Published: (2025)
by: Ge, Chengjie, et al.
Published: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
by: Tu, Xuezhen, et al.
Published: (2026)
by: Tu, Xuezhen, et al.
Published: (2026)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
by: Xu, Wenhao, et al.
Published: (2025)
by: Xu, Wenhao, et al.
Published: (2025)
CAPrompt: Cyclic Prompt Aggregation for Pre-Trained Model Based Class Incremental Learning
by: Li, Qiwei, et al.
Published: (2024)
by: Li, Qiwei, et al.
Published: (2024)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
by: Aparcedo, Alejandro, et al.
Published: (2026)
by: Aparcedo, Alejandro, et al.
Published: (2026)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
by: Li, Peiming, et al.
Published: (2025)
by: Li, Peiming, et al.
Published: (2025)
Spatio-Temporal State Space Model For Efficient Event-Based Optical Flow
by: Humais, Muhammad Ahmed, et al.
Published: (2025)
by: Humais, Muhammad Ahmed, et al.
Published: (2025)
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
by: Liang, Qihua, et al.
Published: (2026)
by: Liang, Qihua, et al.
Published: (2026)
Vectorized Video Representation with Easy Editing via Hierarchical Spatio-Temporally Consistent Proxy Embedding
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
Spatio-temporal Prompting Network for Robust Video Feature Extraction
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
by: Shen, Hao, et al.
Published: (2024)
by: Shen, Hao, et al.
Published: (2024)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
BayesTTA: Continual-Temporal Test-Time Adaptation for Vision-Language Models via Gaussian Discriminant Analysis
by: Cui, Shuang, et al.
Published: (2025)
by: Cui, Shuang, et al.
Published: (2025)
LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
by: Sun, Qihao, et al.
Published: (2026)
by: Sun, Qihao, et al.
Published: (2026)
MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion
by: Xu, Xunnong, et al.
Published: (2024)
by: Xu, Xunnong, et al.
Published: (2024)
Similar Items
-
Token Coordinated Prompt Attention is Needed for Visual Prompting
by: Liu, Zichen, et al.
Published: (2025) -
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
by: Liu, Zichen, et al.
Published: (2025) -
Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
by: Zhou, Jiahuan, et al.
Published: (2025) -
RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
by: Zhu, Kai, et al.
Published: (2026) -
GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision Model
by: Ai, Zixiang, et al.
Published: (2025)