2by2: Weakly-Supervised Learning for Global Action Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bueno-Benito, Elena, Dimiccoli, Mariella
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915068983115776
author Bueno-Benito, Elena
Dimiccoli, Mariella
author_facet Bueno-Benito, Elena
Dimiccoli, Mariella
contents This paper presents a simple yet effective approach for the poorly investigated task of global action segmentation, aiming at grouping frames capturing the same action across videos of different activities. Unlike the case of videos depicting all the same activity, the temporal order of actions is not roughly shared among all videos, making the task even more challenging. We propose to use activity labels to learn, in a weakly-supervised fashion, action representations suitable for global action segmentation. For this purpose, we introduce a triadic learning approach for video pairs, to ensure intra-video action discrimination, as well as inter-video and inter-activity action association. For the backbone architecture, we use a Siamese network based on sparse transformers that takes as input video pairs and determine whether they belong to the same activity. The proposed approach is validated on two challenging benchmark datasets: Breakfast and YouTube Instructions, outperforming state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12829
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 2by2: Weakly-Supervised Learning for Global Action Segmentation
Bueno-Benito, Elena
Dimiccoli, Mariella
Computer Vision and Pattern Recognition
This paper presents a simple yet effective approach for the poorly investigated task of global action segmentation, aiming at grouping frames capturing the same action across videos of different activities. Unlike the case of videos depicting all the same activity, the temporal order of actions is not roughly shared among all videos, making the task even more challenging. We propose to use activity labels to learn, in a weakly-supervised fashion, action representations suitable for global action segmentation. For this purpose, we introduce a triadic learning approach for video pairs, to ensure intra-video action discrimination, as well as inter-video and inter-activity action association. For the backbone architecture, we use a Siamese network based on sparse transformers that takes as input video pairs and determine whether they belong to the same activity. The proposed approach is validated on two challenging benchmark datasets: Breakfast and YouTube Instructions, outperforming state-of-the-art methods.
title 2by2: Weakly-Supervised Learning for Global Action Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.12829