ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Yuchen, Lv, Jianbing, Cheng, Liqi, Meng, Lingyu, Deng, Dazhen, Wu, Yingcai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909621046738944
author He, Yuchen
Lv, Jianbing
Cheng, Liqi
Meng, Lingyu
Deng, Dazhen
Wu, Yingcai
author_facet He, Yuchen
Lv, Jianbing
Cheng, Liqi
Meng, Lingyu
Deng, Dazhen
Wu, Yingcai
contents Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17555
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
He, Yuchen
Lv, Jianbing
Cheng, Liqi
Meng, Lingyu
Deng, Dazhen
Wu, Yingcai
Human-Computer Interaction
Computer Vision and Pattern Recognition
Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework.
title ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17555