AI-Generated Song Detection via Lyrics Transcripts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Frohmann, Markus, Epure, Elena V., Meseguer-Brocal, Gabriel, Schedl, Markus, Hennequin, Romain
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908426603331584
author Frohmann, Markus
Epure, Elena V.
Meseguer-Brocal, Gabriel
Schedl, Markus
Hennequin, Romain
author_facet Frohmann, Markus
Epure, Elena V.
Meseguer-Brocal, Gabriel
Schedl, Markus
Hennequin, Romain
contents The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using audio-based detectors; however, it has been shown that they struggle to generalize to unseen generators or when the audio is perturbed. Furthermore, recent work used accurate and cleanly formatted lyrics sourced from a lyrics provider database to detect AI-generated music. However, in practice, such perfect lyrics are not available (only the audio is); this leaves a substantial gap in applicability in real-life use cases. In this work, we instead propose solving this gap by transcribing songs using general automatic speech recognition (ASR) models. We do this using several detectors. The results on diverse, multi-genre, and multi-lingual lyrics show generally strong detection performance across languages and genres, particularly for our best-performing model using Whisper large-v2 and LLM2Vec embeddings. In addition, we show that our method is more robust than state-of-the-art audio-based ones when the audio is perturbed in different ways and when evaluated on different music generators. Our code is available at https://github.com/deezer/robust-AI-lyrics-detection.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18488
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AI-Generated Song Detection via Lyrics Transcripts
Frohmann, Markus
Epure, Elena V.
Meseguer-Brocal, Gabriel
Schedl, Markus
Hennequin, Romain
Sound
Artificial Intelligence
Computation and Language
The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using audio-based detectors; however, it has been shown that they struggle to generalize to unseen generators or when the audio is perturbed. Furthermore, recent work used accurate and cleanly formatted lyrics sourced from a lyrics provider database to detect AI-generated music. However, in practice, such perfect lyrics are not available (only the audio is); this leaves a substantial gap in applicability in real-life use cases. In this work, we instead propose solving this gap by transcribing songs using general automatic speech recognition (ASR) models. We do this using several detectors. The results on diverse, multi-genre, and multi-lingual lyrics show generally strong detection performance across languages and genres, particularly for our best-performing model using Whisper large-v2 and LLM2Vec embeddings. In addition, we show that our method is more robust than state-of-the-art audio-based ones when the audio is perturbed in different ways and when evaluated on different music generators. Our code is available at https://github.com/deezer/robust-AI-lyrics-detection.
title AI-Generated Song Detection via Lyrics Transcripts
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.18488