STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Jaewoo, Yoon, Jaehong, Kim, Wonjae, Kim, Yunji, Hwang, Sung Ju
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929361274273792
author Lee, Jaewoo
Yoon, Jaehong
Kim, Wonjae
Kim, Yunji
Hwang, Sung Ju
author_facet Lee, Jaewoo
Yoon, Jaehong
Kim, Wonjae
Kim, Yunji
Hwang, Sung Ju
contents Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal correlation between audio-video pairs and multimodal correlation overwriting that forgets audio-video relations. To tackle this problem, we propose a new continual audio-video pre-training method with two novel ideas: (1) Localized Patch Importance Scoring: we introduce a multimodal encoder to determine the importance score for each patch, emphasizing semantically intertwined audio-video patches. (2) Replay-guided Correlation Assessment: to reduce the corruption of previously learned audiovisual knowledge due to drift, we propose to assess the correlation of the current patches on the past steps to identify the patches exhibiting high correlations with the past steps. Based on the results from the two ideas, we perform probabilistic patch selection for effective continual audio-video pre-training. Experimental validation on multiple benchmarks shows that our method achieves a 3.69%p of relative performance gain in zero-shot retrieval tasks compared to strong continual learning baselines, while reducing memory consumption by ~45%.
format Preprint
id arxiv_https___arxiv_org_abs_2310_08204
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
Lee, Jaewoo
Yoon, Jaehong
Kim, Wonjae
Kim, Yunji
Hwang, Sung Ju
Computer Vision and Pattern Recognition
Machine Learning
Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal correlation between audio-video pairs and multimodal correlation overwriting that forgets audio-video relations. To tackle this problem, we propose a new continual audio-video pre-training method with two novel ideas: (1) Localized Patch Importance Scoring: we introduce a multimodal encoder to determine the importance score for each patch, emphasizing semantically intertwined audio-video patches. (2) Replay-guided Correlation Assessment: to reduce the corruption of previously learned audiovisual knowledge due to drift, we propose to assess the correlation of the current patches on the past steps to identify the patches exhibiting high correlations with the past steps. Based on the results from the two ideas, we perform probabilistic patch selection for effective continual audio-video pre-training. Experimental validation on multiple benchmarks shows that our method achieves a 3.69%p of relative performance gain in zero-shot retrieval tasks compared to strong continual learning baselines, while reducing memory consumption by ~45%.
title STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2310.08204