Zero-Shot Temporal Action Localization Through Textual Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liberatori, Benedetta, Conti, Alessandro, Vaquero, Lorenzo, Rota, Paolo, Wang, Yiming, Ricci, Elisa
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917519556608000
author Liberatori, Benedetta
Conti, Alessandro
Vaquero, Lorenzo
Rota, Paolo
Wang, Yiming
Ricci, Elisa
author_facet Liberatori, Benedetta
Conti, Alessandro
Vaquero, Lorenzo
Rota, Paolo
Wang, Yiming
Ricci, Elisa
contents Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), taking advantage of their strong zero-shot transfer capabilities. Yet, these models face evident challenges with fine-grained action classification, making it difficult to directly use them to distinguish between the presence and absence of an action. Most current methods for ZS-TAL address these challenges by training models on large-scale video datasets, which require annotated data and often result in limited generalization performance. Recently, approaches discarding the use of labeled data have emerged as an alternative. Following this direction, we propose a novel approach, ``Textual Guidance for finer localization of actions in videos'' (TEGU), that compensates for the lack of supervision from training data by exploiting rich textual information derived from large language models and structured text extracted from captions. This additional linguistic context can improve fine-grained discrimination by providing richer cues about fine-grained action differences within videos. We validate the effectiveness of the proposed method by conducting experiments on the THUMOS14 and the ActivityNet-v1.3 datasets. Our results show that, by exploiting rich textual information for improved action localization, TEGU outperforms state-of-the-art ZS-TAL approaches that do not involve training
format Preprint
id arxiv_https___arxiv_org_abs_2605_22201
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Shot Temporal Action Localization Through Textual Guidance
Liberatori, Benedetta
Conti, Alessandro
Vaquero, Lorenzo
Rota, Paolo
Wang, Yiming
Ricci, Elisa
Computer Vision and Pattern Recognition
Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), taking advantage of their strong zero-shot transfer capabilities. Yet, these models face evident challenges with fine-grained action classification, making it difficult to directly use them to distinguish between the presence and absence of an action. Most current methods for ZS-TAL address these challenges by training models on large-scale video datasets, which require annotated data and often result in limited generalization performance. Recently, approaches discarding the use of labeled data have emerged as an alternative. Following this direction, we propose a novel approach, ``Textual Guidance for finer localization of actions in videos'' (TEGU), that compensates for the lack of supervision from training data by exploiting rich textual information derived from large language models and structured text extracted from captions. This additional linguistic context can improve fine-grained discrimination by providing richer cues about fine-grained action differences within videos. We validate the effectiveness of the proposed method by conducting experiments on the THUMOS14 and the ActivityNet-v1.3 datasets. Our results show that, by exploiting rich textual information for improved action localization, TEGU outperforms state-of-the-art ZS-TAL approaches that do not involve training
title Zero-Shot Temporal Action Localization Through Textual Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.22201