LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuo, Le, Yuan, Ruibin, Pan, Jiahao, Ma, Yinghao, LI, Yizhi, Zhang, Ge, Liu, Si, Dannenberg, Roger, Fu, Jie, Lin, Chenghua, Benetos, Emmanouil, Xue, Wei, Guo, Yike
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914885682593792
author Zhuo, Le
Yuan, Ruibin
Pan, Jiahao
Ma, Yinghao
LI, Yizhi
Zhang, Ge
Liu, Si
Dannenberg, Roger
Fu, Jie
Lin, Chenghua
Benetos, Emmanouil
Xue, Wei
Guo, Yike
author_facet Zhuo, Le
Yuan, Ruibin
Pan, Jiahao
Ma, Yinghao
LI, Yizhi
Zhang, Ge
Liu, Si
Dannenberg, Roger
Fu, Jie
Lin, Chenghua
Benetos, Emmanouil
Xue, Wei
Guo, Yike
contents We introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as rock and metal. Our novel, training-free approach utilizes Whisper, a weakly supervised robust speech recognition model, and GPT-4, today's most performant chat-based large language model. In the proposed method, Whisper functions as the "ear" by transcribing the audio, while GPT-4 serves as the "brain," acting as an annotator with a strong performance for contextualized output selection and correction. Our experiments show that LyricWhiz significantly reduces Word Error Rate compared to existing methods in English and can effectively transcribe lyrics across multiple languages. Furthermore, we use LyricWhiz to create the first publicly available, large-scale, multilingual lyrics transcription dataset with a CC-BY-NC-SA copyright license, based on MTG-Jamendo, and offer a human-annotated subset for noise level estimation and evaluation. We anticipate that our proposed method and dataset will advance the development of multilingual lyrics transcription, a challenging and emerging task.
format Preprint
id arxiv_https___arxiv_org_abs_2306_17103
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
Zhuo, Le
Yuan, Ruibin
Pan, Jiahao
Ma, Yinghao
LI, Yizhi
Zhang, Ge
Liu, Si
Dannenberg, Roger
Fu, Jie
Lin, Chenghua
Benetos, Emmanouil
Xue, Wei
Guo, Yike
Computation and Language
Sound
Audio and Speech Processing
We introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as rock and metal. Our novel, training-free approach utilizes Whisper, a weakly supervised robust speech recognition model, and GPT-4, today's most performant chat-based large language model. In the proposed method, Whisper functions as the "ear" by transcribing the audio, while GPT-4 serves as the "brain," acting as an annotator with a strong performance for contextualized output selection and correction. Our experiments show that LyricWhiz significantly reduces Word Error Rate compared to existing methods in English and can effectively transcribe lyrics across multiple languages. Furthermore, we use LyricWhiz to create the first publicly available, large-scale, multilingual lyrics transcription dataset with a CC-BY-NC-SA copyright license, based on MTG-Jamendo, and offer a human-annotated subset for noise level estimation and evaluation. We anticipate that our proposed method and dataset will advance the development of multilingual lyrics transcription, a challenging and emerging task.
title LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2306.17103