Introducing Gating and Context into Temporal Action Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Reka, Aglind, Borza, Diana Laura, Reilly, Dominick, Balazia, Michal, Bremond, Francois |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
von: Poddar, Nishit, et al.
Veröffentlicht: (2026)
What Matters in Autonomous Driving Anomaly Detection: A Weakly Supervised Horizon
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024)
Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation Network
von: Sinha, Sanya, et al.
Veröffentlicht: (2025)
von: Sinha, Sanya, et al.
Veröffentlicht: (2025)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets
von: Agrawal, Tanay, et al.
Veröffentlicht: (2025)
von: Agrawal, Tanay, et al.
Veröffentlicht: (2025)
AM Flow: Adapters for Temporal Processing in Action Recognition
von: Agrawal, Tanay, et al.
Veröffentlicht: (2024)
von: Agrawal, Tanay, et al.
Veröffentlicht: (2024)
Temporally Propagated Masks and Bounding Boxes: Combining the Best of Both Worlds for Multi-Object Tracking
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2024)
Are Visual-Language Models Effective in Action Recognition? A Comparative Study
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
von: Ali, Mahmoud, et al.
Veröffentlicht: (2024)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
LAC: Latent Action Composition for Skeleton-based Action Segmentation
von: Yang, Di, et al.
Veröffentlicht: (2023)
von: Yang, Di, et al.
Veröffentlicht: (2023)
No Train Yet Gain: Towards Generic Multi-Object Tracking in Sports and Beyond
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
von: Stanczyk, Tomasz, et al.
Veröffentlicht: (2025)
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
Spatio-Temporal Context Prompting for Zero-Shot Action Detection
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
von: Huang, Wei-Jhe, et al.
Veröffentlicht: (2024)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
von: Kim, Ho-Joong, et al.
Veröffentlicht: (2025)
von: Kim, Ho-Joong, et al.
Veröffentlicht: (2025)
Harnessing Temporal Causality for Advanced Temporal Action Detection
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
von: Cha, Junuk, et al.
Veröffentlicht: (2025)
MVP: Multimodal Emotion Recognition based on Video and Physiological Signals
von: Strizhkova, Valeriya, et al.
Veröffentlicht: (2025)
von: Strizhkova, Valeriya, et al.
Veröffentlicht: (2025)
Prediction-Feedback DETR for Temporal Action Detection
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
Boundary-Recovering Network for Temporal Action Detection
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
Open-Vocabulary Spatio-Temporal Action Detection
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions
von: Zeng, Runhao, et al.
Veröffentlicht: (2024)
von: Zeng, Runhao, et al.
Veröffentlicht: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
von: Qiu, Yicheng, et al.
Veröffentlicht: (2026)
von: Qiu, Yicheng, et al.
Veröffentlicht: (2026)
Anti-Forgetting Adaptation for Unsupervised Person Re-identification
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
Forward Consistency Learning with Gated Context Aggregation for Video Anomaly Detection
von: Lyu, Jiahao, et al.
Veröffentlicht: (2026)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2026)
Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
von: Lachmann, Tim, et al.
Veröffentlicht: (2026)
von: Lachmann, Tim, et al.
Veröffentlicht: (2026)
Dual DETRs for Multi-Label Temporal Action Detection
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2024)
Improving Viewpoint-Invariance and Temporal Consistency for Action Detection
von: Porto, Yannick, et al.
Veröffentlicht: (2026)
von: Porto, Yannick, et al.
Veröffentlicht: (2026)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
von: Ullah, Hayat, et al.
Veröffentlicht: (2025)
von: Ullah, Hayat, et al.
Veröffentlicht: (2025)
LIA-X: Interpretable Latent Portrait Animator
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
CSDN: A Context-Gated Self-Adaptive Detection Network for Real-Time Object Detection
von: Wei, Haolin
Veröffentlicht: (2025)
von: Wei, Haolin
Veröffentlicht: (2025)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
von: Yang, Le, et al.
Veröffentlicht: (2024)
von: Yang, Le, et al.
Veröffentlicht: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
von: Kim, Jihwan, et al.
Veröffentlicht: (2024)
Temporal Action Detection Model Compression by Progressive Block Drop
von: Chen, Xiaoyong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyong, et al.
Veröffentlicht: (2025)
Boundary Discretization and Reliable Classification Network for Temporal Action Detection
von: Fang, Zhenying, et al.
Veröffentlicht: (2023)
von: Fang, Zhenying, et al.
Veröffentlicht: (2023)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
von: Zhu, Xinnan, et al.
Veröffentlicht: (2025)
von: Zhu, Xinnan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
B-MoE: A Body-Part-Aware Mixture-of-Experts "All Parts Matter" Approach to Micro-Action Recognition
von: Poddar, Nishit, et al.
Veröffentlicht: (2026) -
What Matters in Autonomous Driving Anomaly Detection: A Weakly Supervised Horizon
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2024) -
Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation Network
von: Sinha, Sanya, et al.
Veröffentlicht: (2025) -
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025) -
CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction Datasets
von: Agrawal, Tanay, et al.
Veröffentlicht: (2025)