Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
Fuente:
arXiv
Saved in:
| Main Authors: | Gohari, Mahyar, Bestagini, Paolo, Benini, Sergio, Adami, Nicola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
by: Bhagtani, Kratika, et al.
Published: (2024)
by: Bhagtani, Kratika, et al.
Published: (2024)
Multi-View Spectrogram Transformer for Respiratory Sound Classification
by: He, Wentao, et al.
Published: (2023)
by: He, Wentao, et al.
Published: (2023)
Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification
by: Sridhar, Aditya
Published: (2024)
by: Sridhar, Aditya
Published: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
by: Comanducci, Luca, et al.
Published: (2024)
by: Comanducci, Luca, et al.
Published: (2024)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
by: Atito, Sara, et al.
Published: (2022)
by: Atito, Sara, et al.
Published: (2022)
AutoMV: An Automatic Multi-Agent System for Music Video Generation
by: Tang, Xiaoxuan, et al.
Published: (2025)
by: Tang, Xiaoxuan, et al.
Published: (2025)
RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
by: Chen, Jiaben, et al.
Published: (2024)
by: Chen, Jiaben, et al.
Published: (2024)
FLUX that Plays Music
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription
by: Ríos-Vila, Antonio, et al.
Published: (2024)
by: Ríos-Vila, Antonio, et al.
Published: (2024)
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
by: Rinaldi, Ivan, et al.
Published: (2024)
by: Rinaldi, Ivan, et al.
Published: (2024)
Music Genre Classification using Large Language Models
by: Meguenani, Mohamed El Amine, et al.
Published: (2024)
by: Meguenani, Mohamed El Amine, et al.
Published: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks
by: Erattakulangara, Subin, et al.
Published: (2025)
by: Erattakulangara, Subin, et al.
Published: (2025)
Learning Musical Representations for Music Performance Question Answering
by: Diao, Xingjian, et al.
Published: (2025)
by: Diao, Xingjian, et al.
Published: (2025)
FairSSD: Understanding Bias in Synthetic Speech Detectors
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
by: Wahida, Farah, et al.
Published: (2025)
by: Wahida, Farah, et al.
Published: (2025)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
by: Lin, Yan-Bo, et al.
Published: (2024)
by: Lin, Yan-Bo, et al.
Published: (2024)
Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
by: Hisariya, Tanisha, et al.
Published: (2024)
by: Hisariya, Tanisha, et al.
Published: (2024)
End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
MIDGET: Music Conditioned 3D Dance Generation
by: Wang, Jinwu, et al.
Published: (2024)
by: Wang, Jinwu, et al.
Published: (2024)
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
by: Wang, Baisen, et al.
Published: (2024)
by: Wang, Baisen, et al.
Published: (2024)
Learning Sparsity for Effective and Efficient Music Performance Question Answering
by: Diao, Xingjian, et al.
Published: (2025)
by: Diao, Xingjian, et al.
Published: (2025)
Voice Pathology Detection Using Phonation
by: Siva, Sri Raksha, et al.
Published: (2025)
by: Siva, Sri Raksha, et al.
Published: (2025)
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
by: Sun, Luoyi, et al.
Published: (2023)
by: Sun, Luoyi, et al.
Published: (2023)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
by: Li, Ruiqi, et al.
Published: (2024)
by: Li, Ruiqi, et al.
Published: (2024)
DanceChat: Large Language Model-Guided Music-to-Dance Generation
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
FilmComposer: LLM-Driven Music Production for Silent Film Clips
by: Xie, Zhifeng, et al.
Published: (2025)
by: Xie, Zhifeng, et al.
Published: (2025)
GCDance: Genre-Controlled Music-Driven 3D Full Body Dance Generation
by: Liu, Xinran, et al.
Published: (2025)
by: Liu, Xinran, et al.
Published: (2025)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
by: Klein, Nicholas, et al.
Published: (2025)
by: Klein, Nicholas, et al.
Published: (2025)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
by: Ghosh, Anindita, et al.
Published: (2025)
by: Ghosh, Anindita, et al.
Published: (2025)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
by: Dai, Yuqin, et al.
Published: (2025)
by: Dai, Yuqin, et al.
Published: (2025)
Enhancing CTC-Based Visual Speech Recognition
by: Laux, Hendrik, et al.
Published: (2024)
by: Laux, Hendrik, et al.
Published: (2024)
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
by: Huang, Mingyang, et al.
Published: (2025)
by: Huang, Mingyang, et al.
Published: (2025)
A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
by: Ma, Junwen, et al.
Published: (2026)
by: Ma, Junwen, et al.
Published: (2026)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
by: D., Quang-Anh N., et al.
Published: (2024)
by: D., Quang-Anh N., et al.
Published: (2024)
Similar Items
-
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024) -
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
by: Bhagtani, Kratika, et al.
Published: (2024) -
Multi-View Spectrogram Transformer for Respiratory Sound Classification
by: He, Wentao, et al.
Published: (2023) -
Attention-guided Spectrogram Sequence Modeling with CNNs for Music Genre Classification
by: Sridhar, Aditya
Published: (2024) -
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
by: Comanducci, Luca, et al.
Published: (2024)