One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Trung Thanh, Kawanishi, Yasutomo, Komamizu, Takahiro, Ide, Ichiro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
by: Liu, Tingwei, et al.
Published: (2024)
by: Liu, Tingwei, et al.
Published: (2024)
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
by: Chen, Junan, et al.
Published: (2025)
by: Chen, Junan, et al.
Published: (2025)
ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
by: Nguyen, Trung Thanh, et al.
Published: (2026)
by: Nguyen, Trung Thanh, et al.
Published: (2026)
Small Object Detection for Birds with Swin Transformer
by: Huo, Da, et al.
Published: (2025)
by: Huo, Da, et al.
Published: (2025)
Multi-View Video-Based Learning: Leveraging Weak Labels for Frame-Level Perception
by: John, Vijay, et al.
Published: (2024)
by: John, Vijay, et al.
Published: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
by: Zhang, Bolin, et al.
Published: (2026)
by: Zhang, Bolin, et al.
Published: (2026)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
by: Hu, Zhanjie, et al.
Published: (2026)
by: Hu, Zhanjie, et al.
Published: (2026)
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
by: Fang, Zhenying, et al.
Published: (2025)
by: Fang, Zhenying, et al.
Published: (2025)
Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements
by: Kinoshita, Genki, et al.
Published: (2026)
by: Kinoshita, Genki, et al.
Published: (2026)
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025)
by: Sia, Zhen Hao, et al.
Published: (2025)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
Full-Stage Pseudo Label Quality Enhancement for Weakly-supervised Temporal Action Localization
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
Harnessing Temporal Causality for Advanced Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2024)
by: Liu, Shuming, et al.
Published: (2024)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
by: Ullah, Hayat, et al.
Published: (2025)
by: Ullah, Hayat, et al.
Published: (2025)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
Density-Guided Label Smoothing for Temporal Localization of Driving Actions
by: Alkanat, Tunc, et al.
Published: (2024)
by: Alkanat, Tunc, et al.
Published: (2024)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions
by: Zeng, Runhao, et al.
Published: (2024)
by: Zeng, Runhao, et al.
Published: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
by: Qiu, Yicheng, et al.
Published: (2026)
by: Qiu, Yicheng, et al.
Published: (2026)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
Prediction-Feedback DETR for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
by: Gowda, Shreyank N, et al.
Published: (2025)
by: Gowda, Shreyank N, et al.
Published: (2025)
Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos
by: Ochin, Jeremie, et al.
Published: (2025)
by: Ochin, Jeremie, et al.
Published: (2025)
Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
by: Duong, Anh-Kiet, et al.
Published: (2025)
by: Duong, Anh-Kiet, et al.
Published: (2025)
Improving Viewpoint-Invariance and Temporal Consistency for Action Detection
by: Porto, Yannick, et al.
Published: (2026)
by: Porto, Yannick, et al.
Published: (2026)
DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
by: Cheng, Haozhe, et al.
Published: (2024)
by: Cheng, Haozhe, et al.
Published: (2024)
POTLoc: Pseudo-Label Oriented Transformer for Point-Supervised Temporal Action Localization
by: Vahdani, Elahe, et al.
Published: (2023)
by: Vahdani, Elahe, et al.
Published: (2023)
Similar Items
-
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024) -
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
by: Liu, Tingwei, et al.
Published: (2024)