A Survey of Foundation Models for Music Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Wenjun, Cai, Ying, Wu, Ziyang, Zhang, Wenyi, Chen, Yifan, Qi, Rundong, Dong, Mengqi, Chen, Peigen, Dong, Xiao, Shi, Fenghao, Guo, Lei, Han, Junwei, Ge, Bao, Liu, Tianming, Gan, Lin, Zhang, Tuo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913500346974208
author Li, Wenjun
Cai, Ying
Wu, Ziyang
Zhang, Wenyi
Chen, Yifan
Qi, Rundong
Dong, Mengqi
Chen, Peigen
Dong, Xiao
Shi, Fenghao
Guo, Lei
Han, Junwei
Ge, Bao
Liu, Tianming
Gan, Lin
Zhang, Tuo
author_facet Li, Wenjun
Cai, Ying
Wu, Ziyang
Zhang, Wenyi
Chen, Yifan
Qi, Rundong
Dong, Mengqi
Chen, Peigen
Dong, Xiao
Shi, Fenghao
Guo, Lei
Han, Junwei
Ge, Bao
Liu, Tianming
Gan, Lin
Zhang, Tuo
contents Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections. The rapid advancement of artificial intelligence (AI) has introduced new ways to analyze music, aiming to replicate human understanding of music and provide related services. While the traditional models focused on audio features and simple tasks, the recent development of large language models (LLMs) and foundation models (FMs), which excel in various fields by integrating semantic information and demonstrating strong reasoning abilities, could capture complex musical features and patterns, integrate music with language and incorporate rich musical, emotional and psychological knowledge. Therefore, they have the potential in handling complex music understanding tasks from a semantic perspective, producing outputs closer to human perception. This work, to our best knowledge, is one of the early reviews of the intersection of AI techniques and music understanding. We investigated, analyzed, and tested recent large-scale music foundation models in respect of their music comprehension abilities. We also discussed their limitations and proposed possible future directions, offering insights for researchers in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09601
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey of Foundation Models for Music Understanding
Li, Wenjun
Cai, Ying
Wu, Ziyang
Zhang, Wenyi
Chen, Yifan
Qi, Rundong
Dong, Mengqi
Chen, Peigen
Dong, Xiao
Shi, Fenghao
Guo, Lei
Han, Junwei
Ge, Bao
Liu, Tianming
Gan, Lin
Zhang, Tuo
Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections. The rapid advancement of artificial intelligence (AI) has introduced new ways to analyze music, aiming to replicate human understanding of music and provide related services. While the traditional models focused on audio features and simple tasks, the recent development of large language models (LLMs) and foundation models (FMs), which excel in various fields by integrating semantic information and demonstrating strong reasoning abilities, could capture complex musical features and patterns, integrate music with language and incorporate rich musical, emotional and psychological knowledge. Therefore, they have the potential in handling complex music understanding tasks from a semantic perspective, producing outputs closer to human perception. This work, to our best knowledge, is one of the early reviews of the intersection of AI techniques and music understanding. We investigated, analyzed, and tested recent large-scale music foundation models in respect of their music comprehension abilities. We also discussed their limitations and proposed possible future directions, offering insights for researchers in this field.
title A Survey of Foundation Models for Music Understanding
topic Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2409.09601