Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
Fuente:
arXiv
Saved in:
| Main Authors: | Frohmann, Markus, Meseguer-Brocal, Gabriel, Schedl, Markus, Epure, Elena V. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Generated Song Detection via Lyrics Transcripts
by: Frohmann, Markus, et al.
Published: (2025)
by: Frohmann, Markus, et al.
Published: (2025)
AI-Generated Music Detection and its Challenges
by: Afchar, Darius, et al.
Published: (2025)
by: Afchar, Darius, et al.
Published: (2025)
S-KEY: Self-supervised Learning of Major and Minor Keys from Audio
by: Kong, Yuexuan, et al.
Published: (2025)
by: Kong, Yuexuan, et al.
Published: (2025)
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
by: Torres, Bernardo, et al.
Published: (2025)
by: Torres, Bernardo, et al.
Published: (2025)
Detecting music deepfakes is easy but actually hard
by: Afchar, Darius, et al.
Published: (2024)
by: Afchar, Darius, et al.
Published: (2024)
An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging
by: Meseguer-Brocal, Gabriel, et al.
Published: (2024)
by: Meseguer-Brocal, Gabriel, et al.
Published: (2024)
STONE: Self-supervised Tonality Estimator
by: Kong, Yuexuan, et al.
Published: (2024)
by: Kong, Yuexuan, et al.
Published: (2024)
From Real to Cloned Singer Identification
by: Desblancs, Dorian, et al.
Published: (2024)
by: Desblancs, Dorian, et al.
Published: (2024)
Computational Narrative Understanding for Expressive Text-to-Speech
by: Michel, Gaspard, et al.
Published: (2025)
by: Michel, Gaspard, et al.
Published: (2025)
Synthetic Lyrics Detection Across Languages and Genres
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Emergent musical properties of a transformer under contrastive self-supervised learning
by: Kong, Yuexuan, et al.
Published: (2025)
by: Kong, Yuexuan, et al.
Published: (2025)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
by: Zhuo, Le, et al.
Published: (2023)
by: Zhuo, Le, et al.
Published: (2023)
Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
by: Fuglsig, Andreas Jonas, et al.
Published: (2026)
by: Fuglsig, Andreas Jonas, et al.
Published: (2026)
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
by: Li, Yupei, et al.
Published: (2024)
by: Li, Yupei, et al.
Published: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
by: Wilkins, Julia, et al.
Published: (2024)
by: Wilkins, Julia, et al.
Published: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
by: Jin, Zhan, et al.
Published: (2025)
by: Jin, Zhan, et al.
Published: (2025)
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
by: Huang, Jiawen, et al.
Published: (2025)
by: Huang, Jiawen, et al.
Published: (2025)
Perturbed Public Voices (P$^{2}$V): A Dataset for Robust Audio Deepfake Detection
by: Gao, Chongyang, et al.
Published: (2025)
by: Gao, Chongyang, et al.
Published: (2025)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
by: Syed, Jaza, et al.
Published: (2025)
by: Syed, Jaza, et al.
Published: (2025)
An Octave-based Multi-Resolution CQT Architecture for Diffusion-based Audio Generation
by: da Costa, Maurício do V. M., et al.
Published: (2025)
by: da Costa, Maurício do V. M., et al.
Published: (2025)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Towards Robust Audio Deepfake Detection: A Evolving Benchmark for Continual Learning
by: Zhang, Xiaohui, et al.
Published: (2024)
by: Zhang, Xiaohui, et al.
Published: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
by: Meng, Hao, et al.
Published: (2026)
by: Meng, Hao, et al.
Published: (2026)
LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
by: Shekar, Ram C. M. C., et al.
Published: (2025)
by: Shekar, Ram C. M. C., et al.
Published: (2025)
Scalable Music Cover Retrieval Using Lyrics-Aligned Audio Embeddings
by: Affolter, Joanne, et al.
Published: (2026)
by: Affolter, Joanne, et al.
Published: (2026)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
by: Kheir, Yassine El, et al.
Published: (2026)
by: Kheir, Yassine El, et al.
Published: (2026)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
by: Araz, R. Oguz, et al.
Published: (2025)
by: Araz, R. Oguz, et al.
Published: (2025)
SongCreator: Lyrics-based Universal Song Generation
by: Lei, Shun, et al.
Published: (2024)
by: Lei, Shun, et al.
Published: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
by: Wang, Zehan, et al.
Published: (2025)
by: Wang, Zehan, et al.
Published: (2025)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
by: Niu, Zhikang, et al.
Published: (2024)
by: Niu, Zhikang, et al.
Published: (2024)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Robust Lossy Audio Compression Identification
by: Koops, Hendrik Vincent, et al.
Published: (2024)
by: Koops, Hendrik Vincent, et al.
Published: (2024)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
by: Han, Bing, et al.
Published: (2025)
by: Han, Bing, et al.
Published: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
by: Yang, Dongchao, et al.
Published: (2023)
by: Yang, Dongchao, et al.
Published: (2023)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
by: Khanjani, Zahra, et al.
Published: (2024)
by: Khanjani, Zahra, et al.
Published: (2024)
A Noval Feature via Color Quantisation for Fake Audio Detection
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis
by: Zou, Dinghao, et al.
Published: (2025)
by: Zou, Dinghao, et al.
Published: (2025)
Similar Items
-
AI-Generated Song Detection via Lyrics Transcripts
by: Frohmann, Markus, et al.
Published: (2025) -
AI-Generated Music Detection and its Challenges
by: Afchar, Darius, et al.
Published: (2025) -
S-KEY: Self-supervised Learning of Major and Minor Keys from Audio
by: Kong, Yuexuan, et al.
Published: (2025) -
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
by: Torres, Bernardo, et al.
Published: (2025) -
Detecting music deepfakes is easy but actually hard
by: Afchar, Darius, et al.
Published: (2024)