ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909621046738944 |
|---|---|
| author | He, Yuchen Lv, Jianbing Cheng, Liqi Meng, Lingyu Deng, Dazhen Wu, Yingcai |
| author_facet | He, Yuchen Lv, Jianbing Cheng, Liqi Meng, Lingyu Deng, Dazhen Wu, Yingcai |
| contents | Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17555 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization He, Yuchen Lv, Jianbing Cheng, Liqi Meng, Lingyu Deng, Dazhen Wu, Yingcai Human-Computer Interaction Computer Vision and Pattern Recognition Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework. |
| title | ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization |
| topic | Human-Computer Interaction Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.17555 |