Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Duong, Anh-Kiet, Gomez-Krämer, Petra
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915670987374592
author Duong, Anh-Kiet
Gomez-Krämer, Petra
author_facet Duong, Anh-Kiet
Gomez-Krämer, Petra
contents We present our solution to the BinEgo-360 Challenge at ICCV 2025, which focuses on temporal action localization (TAL) in multi-perspective and multi-modal video settings. The challenge provides a dataset containing panoramic, third-person, and egocentric recordings, annotated with fine-grained action classes. Our approach is built on the Temporal Shift Module (TSM), which we extend to handle TAL by introducing a background class and classifying fixed-length non-overlapping intervals. We employ a multi-task learning framework that jointly optimizes for scene classification and TAL, leveraging contextual cues between actions and environments. Finally, we integrate multiple models through a weighted ensemble strategy, which improves robustness and consistency of predictions. Our method is ranked first in both the initial and extended rounds of the competition, demonstrating the effectiveness of combining multi-task learning, an efficient backbone, and ensemble learning for TAL.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
Duong, Anh-Kiet
Gomez-Krämer, Petra
Computer Vision and Pattern Recognition
We present our solution to the BinEgo-360 Challenge at ICCV 2025, which focuses on temporal action localization (TAL) in multi-perspective and multi-modal video settings. The challenge provides a dataset containing panoramic, third-person, and egocentric recordings, annotated with fine-grained action classes. Our approach is built on the Temporal Shift Module (TSM), which we extend to handle TAL by introducing a background class and classifying fixed-length non-overlapping intervals. We employ a multi-task learning framework that jointly optimizes for scene classification and TAL, leveraging contextual cues between actions and environments. Finally, we integrate multiple models through a weighted ensemble strategy, which improves robustness and consistency of predictions. Our method is ranked first in both the initial and extended rounds of the competition, demonstrating the effectiveness of combining multi-task learning, an efficient backbone, and ensemble learning for TAL.
title Multi-task Learning with Extended Temporal Shift Module for Temporal Action Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11189