Multi Activity Sequence Alignment via Implicit Clustering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kwon, Taein, Pataki, Zador, Rad, Mahdi, Pollefeys, Marc
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929761775779840
author Kwon, Taein
Pataki, Zador
Rad, Mahdi
Pollefeys, Marc
author_facet Kwon, Taein
Pataki, Zador
Rad, Mahdi
Pollefeys, Marc
contents Self-supervised temporal sequence alignment can provide rich and effective representations for a wide range of applications. However, existing methods for achieving optimal performance are mostly limited to aligning sequences of the same activity only and require separate models to be trained for each activity. We propose a novel framework that overcomes these limitations using sequence alignment via implicit clustering. Specifically, our key idea is to perform implicit clip-level clustering while aligning frames in sequences. This coupled with our proposed dual augmentation technique enhances the network's ability to learn generalizable and discriminative representations. Our experiments show that our proposed method outperforms state-of-the-art results and highlight the generalization capability of our framework with multi activity and different modalities on three diverse datasets, H2O, PennAction, and IKEA ASM. We will release our code upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12519
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi Activity Sequence Alignment via Implicit Clustering
Kwon, Taein
Pataki, Zador
Rad, Mahdi
Pollefeys, Marc
Computer Vision and Pattern Recognition
Self-supervised temporal sequence alignment can provide rich and effective representations for a wide range of applications. However, existing methods for achieving optimal performance are mostly limited to aligning sequences of the same activity only and require separate models to be trained for each activity. We propose a novel framework that overcomes these limitations using sequence alignment via implicit clustering. Specifically, our key idea is to perform implicit clip-level clustering while aligning frames in sequences. This coupled with our proposed dual augmentation technique enhances the network's ability to learn generalizable and discriminative representations. Our experiments show that our proposed method outperforms state-of-the-art results and highlight the generalization capability of our framework with multi activity and different modalities on three diverse datasets, H2O, PennAction, and IKEA ASM. We will release our code upon acceptance.
title Multi Activity Sequence Alignment via Implicit Clustering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12519