Open-Vocabulary Spatio-Temporal Action Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Tao, Ge, Shuqiu, Qin, Jie, Wu, Gangshan, Wang, Limin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
STMixer: A One-Stage Sparse Action Detector
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)
by: Cui, Yutao, et al.
Published: (2023)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
by: Luo, Xingyu, et al.
Published: (2026)
by: Luo, Xingyu, et al.
Published: (2026)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025)
by: Sia, Zhen Hao, et al.
Published: (2025)
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
Spatial-Temporal Human-Object Interaction Detection
by: Sun, Xu, et al.
Published: (2025)
by: Sun, Xu, et al.
Published: (2025)
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
by: Chao, Yuhao, et al.
Published: (2025)
by: Chao, Yuhao, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Learning to Detect and Segment for Open Vocabulary Object Detection
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Pure-Pass: Fine-Grained, Adaptive Masking for Dynamic Token-Mixing Routing in Lightweight Image Super-Resolution
by: Wu, Junyu, et al.
Published: (2025)
by: Wu, Junyu, et al.
Published: (2025)
Sketch and Refine: Towards Fast and Accurate Lane Detection
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
Open-Vocabulary Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
Evaluating the Performance of Open-Vocabulary Object Detection in Low-quality Image
by: Wu, Po-Chih
Published: (2025)
by: Wu, Po-Chih
Published: (2025)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
by: Zhao, Zhiyu, et al.
Published: (2023)
by: Zhao, Zhiyu, et al.
Published: (2023)
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
by: Cheng, Haozhe, et al.
Published: (2024)
by: Cheng, Haozhe, et al.
Published: (2024)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
by: Fang, Zhenying, et al.
Published: (2025)
by: Fang, Zhenying, et al.
Published: (2025)
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025)
by: Zheng, Haozhen, et al.
Published: (2025)
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
by: Liu, Haijing, et al.
Published: (2025)
by: Liu, Haijing, et al.
Published: (2025)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
by: Patel, Naman, et al.
Published: (2025)
by: Patel, Naman, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Efficient Test-Time Prompt Tuning for Vision-Language Models
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
by: Han, Tingting, et al.
Published: (2026)
by: Han, Tingting, et al.
Published: (2026)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
by: Yuan, Zhenlong, et al.
Published: (2025)
by: Yuan, Zhenlong, et al.
Published: (2025)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
by: Huang, Wei-Jhe, et al.
Published: (2024)
by: Huang, Wei-Jhe, et al.
Published: (2024)
Similar Items
-
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024) -
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024) -
STMixer: A One-Stage Sparse Action Detector
by: Wu, Tao, et al.
Published: (2024) -
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026) -
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)