TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Ho-Joong, Hong, Jung-Ho, Kong, Heejo, Lee, Seong-Whan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909160337047552
author Kim, Ho-Joong
Hong, Jung-Ho
Kong, Heejo
Lee, Seong-Whan
author_facet Kim, Ho-Joong
Hong, Jung-Ho
Kong, Heejo
Lee, Seong-Whan
contents In this paper, we investigate that the normalized coordinate expression is a key factor as reliance on hand-crafted components in query-based detectors for temporal action detection (TAD). Despite significant advancements towards an end-to-end framework in object detection, query-based detectors have been limited in achieving full end-to-end modeling in TAD. To address this issue, we propose \modelname{}, a full end-to-end temporal action detection transformer that integrates time-aligned coordinate expression. We reformulate coordinate expression utilizing actual timeline values, ensuring length-invariant representations from the extremely diverse video duration environment. Furthermore, our proposed adaptive query selection dynamically adjusts the number of queries based on video length, providing a suitable solution for varying video durations compared to a fixed query set. Our approach not only simplifies the TAD process by eliminating the need for hand-crafted components but also significantly improves the performance of query-based detectors. Our TE-TAD outperforms the previous query-based detectors and achieves competitive performance compared to state-of-the-art methods on popular benchmark datasets. Code is available at: https://github.com/Dotori-HJ/TE-TAD
format Preprint
id arxiv_https___arxiv_org_abs_2404_02405
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
Kim, Ho-Joong
Hong, Jung-Ho
Kong, Heejo
Lee, Seong-Whan
Computer Vision and Pattern Recognition
In this paper, we investigate that the normalized coordinate expression is a key factor as reliance on hand-crafted components in query-based detectors for temporal action detection (TAD). Despite significant advancements towards an end-to-end framework in object detection, query-based detectors have been limited in achieving full end-to-end modeling in TAD. To address this issue, we propose \modelname{}, a full end-to-end temporal action detection transformer that integrates time-aligned coordinate expression. We reformulate coordinate expression utilizing actual timeline values, ensuring length-invariant representations from the extremely diverse video duration environment. Furthermore, our proposed adaptive query selection dynamically adjusts the number of queries based on video length, providing a suitable solution for varying video durations compared to a fixed query set. Our approach not only simplifies the TAD process by eliminating the need for hand-crafted components but also significantly improves the performance of query-based detectors. Our TE-TAD outperforms the previous query-based detectors and achieves competitive performance compared to state-of-the-art methods on popular benchmark datasets. Code is available at: https://github.com/Dotori-HJ/TE-TAD
title TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.02405