Segment Transformer: AI-Generated Music Detection via Music Structural Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yumin, Go, Seonghyeon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915488378912768
author Kim, Yumin
Go, Seonghyeon
author_facet Kim, Yumin
Go, Seonghyeon
contents Audio and music generation systems have been remarkably developed in the music information retrieval (MIR) research field. The advancement of these technologies raises copyright concerns, as ownership and authorship of AI-generated music (AIGM) remain unclear. Also, it can be difficult to determine whether a piece was generated by AI or composed by humans clearly. To address these challenges, we aim to improve the accuracy of AIGM detection by analyzing the structural patterns of music segments. Specifically, to extract musical features from short audio clips, we integrated various pre-trained models, including self-supervised learning (SSL) models or an audio effect encoder, each within our suggested transformer-based framework. Furthermore, for long audio, we developed a segment transformer that divides music into segments and learns inter-segment relationships. We used the FakeMusicCaps and SONICS datasets, achieving high accuracy in both the short-audio and full-audio detection experiments. These findings suggest that integrating segment-level musical features into long-range temporal analysis can effectively enhance both the performance and robustness of AIGM detection systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
Kim, Yumin
Go, Seonghyeon
Sound
Artificial Intelligence
Audio and Speech Processing
Audio and music generation systems have been remarkably developed in the music information retrieval (MIR) research field. The advancement of these technologies raises copyright concerns, as ownership and authorship of AI-generated music (AIGM) remain unclear. Also, it can be difficult to determine whether a piece was generated by AI or composed by humans clearly. To address these challenges, we aim to improve the accuracy of AIGM detection by analyzing the structural patterns of music segments. Specifically, to extract musical features from short audio clips, we integrated various pre-trained models, including self-supervised learning (SSL) models or an audio effect encoder, each within our suggested transformer-based framework. Furthermore, for long audio, we developed a segment transformer that divides music into segments and learns inter-segment relationships. We used the FakeMusicCaps and SONICS datasets, achieving high accuracy in both the short-audio and full-audio detection experiments. These findings suggest that integrating segment-level musical features into long-range temporal analysis can effectively enhance both the performance and robustness of AIGM detection systems.
title Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.08283