A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vilaca, Luis, Yu, Yi, Vinan, Paula
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917852790915072
author Vilaca, Luis
Yu, Yi
Vinan, Paula
author_facet Vilaca, Luis
Yu, Yi
Vinan, Paula
contents Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey. Besides analyzing the models used in this context, we also discuss some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate objective functions frequently used and discuss how audio-visual data is exploited in the optimization process, i.e., the different methodologies for representing knowledge in the audio-visual domain. In fact, we focus on how human-understandable mechanisms, i.e., structured knowledge that reflects comprehensible knowledge, can guide the learning process. Most importantly, we provide a summarization of the recent progress of Audio-Visual Correlation Learning (AVCL) and discuss the future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
Vilaca, Luis
Yu, Yi
Vinan, Paula
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey. Besides analyzing the models used in this context, we also discuss some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate objective functions frequently used and discuss how audio-visual data is exploited in the optimization process, i.e., the different methodologies for representing knowledge in the audio-visual domain. In fact, we focus on how human-understandable mechanisms, i.e., structured knowledge that reflects comprehensible knowledge, can guide the learning process. Most importantly, we provide a summarization of the recent progress of Audio-Visual Correlation Learning (AVCL) and discuss the future research directions.
title A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.00049