STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zichen, Xu, Kunlun, Su, Bing, Zou, Xu, Peng, Yuxin, Zhou, Jiahuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
Token Coordinated Prompt Attention is Needed for Visual Prompting
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Componential Prompt-Knowledge Alignment for Domain Incremental Learning
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification
von: Xu, Kunlun, et al.
Veröffentlicht: (2024)
von: Xu, Kunlun, et al.
Veröffentlicht: (2024)
Selective Visual Prompting in Vision Mamba
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
von: Yao, Yifeng, et al.
Veröffentlicht: (2024)
Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-Identification
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
von: Xu, Kunlun, et al.
Veröffentlicht: (2026)
von: Xu, Kunlun, et al.
Veröffentlicht: (2026)
Vision Graph Prompting via Semantic Low-Rank Decomposition
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
GAPrompt: Geometry-Aware Point Cloud Prompt for 3D Vision Model
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
C^2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
von: Xu, Kunlun, et al.
Veröffentlicht: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
UPP: Unified Point-Level Prompting for Robust Point Cloud Analysis
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
von: Ai, Zixiang, et al.
Veröffentlicht: (2025)
Class-aware Domain Knowledge Fusion and Fission for Continual Test-Time Adaptation
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
von: Yang, Min, et al.
Veröffentlicht: (2024)
von: Yang, Min, et al.
Veröffentlicht: (2024)
Bi-C2R: Bidirectional Continual Compatible Representation for Re-indexing Free Lifelong Person Re-identification
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
CKDA: Cross-modality Knowledge Disentanglement and Alignment for Visible-Infrared Lifelong Person Re-identification
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
von: Cui, Zhenyu, et al.
Veröffentlicht: (2025)
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2025)
Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition
von: Xu, Xiaogang, et al.
Veröffentlicht: (2024)
von: Xu, Xiaogang, et al.
Veröffentlicht: (2024)
Low-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm
von: Xu, Xiaogang, et al.
Veröffentlicht: (2026)
von: Xu, Xiaogang, et al.
Veröffentlicht: (2026)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
von: Shen, Cuifeng, et al.
Veröffentlicht: (2025)
von: Shen, Cuifeng, et al.
Veröffentlicht: (2025)
CAPrompt: Cyclic Prompt Aggregation for Pre-Trained Model Based Class Incremental Learning
von: Li, Qiwei, et al.
Veröffentlicht: (2024)
von: Li, Qiwei, et al.
Veröffentlicht: (2024)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
von: Xu, Zhiyang, et al.
Veröffentlicht: (2026)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2026)
CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
von: Su, Rui, et al.
Veröffentlicht: (2019)
von: Su, Rui, et al.
Veröffentlicht: (2019)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
von: Li, Geng, et al.
Veröffentlicht: (2025)
von: Li, Geng, et al.
Veröffentlicht: (2025)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models
von: Xu, Jinglin, et al.
Veröffentlicht: (2024)
von: Xu, Jinglin, et al.
Veröffentlicht: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
von: Khare, Avishree, et al.
Veröffentlicht: (2025)
von: Khare, Avishree, et al.
Veröffentlicht: (2025)
Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
von: Zhou, Shengchao, et al.
Veröffentlicht: (2025)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
Incremental Object Keypoint Learning
von: Liang, Mingfu, et al.
Veröffentlicht: (2025)
von: Liang, Mingfu, et al.
Veröffentlicht: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
von: Zhao, Henghao, et al.
Veröffentlicht: (2025)
von: Zhao, Henghao, et al.
Veröffentlicht: (2025)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025) -
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025) -
Token Coordinated Prompt Attention is Needed for Visual Prompting
von: Liu, Zichen, et al.
Veröffentlicht: (2025) -
Componential Prompt-Knowledge Alignment for Domain Incremental Learning
von: Xu, Kunlun, et al.
Veröffentlicht: (2025) -
DASK: Distribution Rehearsing via Adaptive Style Kernel Learning for Exemplar-Free Lifelong Person Re-Identification
von: Xu, Kunlun, et al.
Veröffentlicht: (2024)