CVSM: Contrastive Vocal Similarity Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Garoufis, Christos, Zlatintsi, Athanasia, Maragos, Petros |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-training Music Classification Models via Music Source Separation
by: Garoufis, Christos, et al.
Published: (2023)
by: Garoufis, Christos, et al.
Published: (2023)
Classical Guitar Duet Separation using GuitarDuets -- a Dataset of Real and Synthesized Guitar Recordings
by: Glytsos, Marios, et al.
Published: (2025)
by: Glytsos, Marios, et al.
Published: (2025)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
by: Wang, Ju-Chiang, et al.
Published: (2024)
by: Wang, Ju-Chiang, et al.
Published: (2024)
voc2vec: A Foundation Model for Non-Verbal Vocalization
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
by: Wei, Haojie, et al.
Published: (2023)
by: Wei, Haojie, et al.
Published: (2023)
DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions
by: Yu, Chin-Yun, et al.
Published: (2025)
by: Yu, Chin-Yun, et al.
Published: (2025)
Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
by: Lu, Minhui, et al.
Published: (2026)
by: Lu, Minhui, et al.
Published: (2026)
Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
by: Bhake, Yash, et al.
Published: (2025)
by: Bhake, Yash, et al.
Published: (2025)
Auditory Representation Effective for Estimating Vocal Tract Information
by: Irino, Toshio, et al.
Published: (2023)
by: Irino, Toshio, et al.
Published: (2023)
DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
by: Wei, Haojie, et al.
Published: (2024)
by: Wei, Haojie, et al.
Published: (2024)
A Reliable and Efficient Detection Pipeline for Rodent Ultrasonic Vocalizations
by: Anis, Sabah Shahnoor, et al.
Published: (2025)
by: Anis, Sabah Shahnoor, et al.
Published: (2025)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
by: Nobukawa, Rinka, et al.
Published: (2025)
by: Nobukawa, Rinka, et al.
Published: (2025)
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
by: Miron, Marius, et al.
Published: (2024)
by: Miron, Marius, et al.
Published: (2024)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
by: Li, Jialu, et al.
Published: (2023)
by: Li, Jialu, et al.
Published: (2023)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
by: Kheir, Yassine El, et al.
Published: (2024)
by: Kheir, Yassine El, et al.
Published: (2024)
Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
by: Chen, Yu-Wen, et al.
Published: (2025)
by: Chen, Yu-Wen, et al.
Published: (2025)
VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Computational Extraction of Intonation and Tuning Systems from Multiple Microtonal Monophonic Vocal Recordings with Diverse Modes
by: Shafiei, Sepideh, et al.
Published: (2025)
by: Shafiei, Sepideh, et al.
Published: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
by: Zhu, Ge, et al.
Published: (2024)
by: Zhu, Ge, et al.
Published: (2024)
Domain Adaptation for Contrastive Audio-Language Models
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack
by: Liu, Yuxuan, et al.
Published: (2025)
by: Liu, Yuxuan, et al.
Published: (2025)
Live Vocal Extraction from K-pop Performances
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
by: Maghsoudi, Maryam, et al.
Published: (2026)
by: Maghsoudi, Maryam, et al.
Published: (2026)
Learning Separated Representations for Instrument-based Music Similarity
by: Hashizume, Yuka, et al.
Published: (2025)
by: Hashizume, Yuka, et al.
Published: (2025)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
by: Tian, Haokun, et al.
Published: (2025)
by: Tian, Haokun, et al.
Published: (2025)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
by: Salimi, Amir, et al.
Published: (2025)
by: Salimi, Amir, et al.
Published: (2025)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
by: Borodin, Kirill, et al.
Published: (2026)
by: Borodin, Kirill, et al.
Published: (2026)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
by: Hashizume, Yuka, et al.
Published: (2024)
by: Hashizume, Yuka, et al.
Published: (2024)
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
by: Shi, Runwu, et al.
Published: (2024)
by: Shi, Runwu, et al.
Published: (2024)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
by: Panah, Davoud Shariat, et al.
Published: (2025)
by: Panah, Davoud Shariat, et al.
Published: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
by: Jing, Xin, et al.
Published: (2024)
by: Jing, Xin, et al.
Published: (2024)
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
by: Imamura, Takehiro, et al.
Published: (2025)
by: Imamura, Takehiro, et al.
Published: (2025)
GESI: Gammachirp Envelope Similarity Index for Predicting Intelligibility of Simulated Hearing Loss Sounds
by: Yamamoto, Ayako, et al.
Published: (2023)
by: Yamamoto, Ayako, et al.
Published: (2023)
SCOREQ: Speech Quality Assessment with Contrastive Regression
by: Ragano, Alessandro, et al.
Published: (2024)
by: Ragano, Alessandro, et al.
Published: (2024)
Noise-Aware Speech Separation with Contrastive Learning
by: Zhang, Zizheng, et al.
Published: (2023)
by: Zhang, Zizheng, et al.
Published: (2023)
Speaker Contrastive Learning for Source Speaker Tracing
by: Wang, Qing, et al.
Published: (2024)
by: Wang, Qing, et al.
Published: (2024)
Towards the Synthesis of Non-speech Vocalizations
by: Hoq, Enjamamul, et al.
Published: (2024)
by: Hoq, Enjamamul, et al.
Published: (2024)
Similar Items
-
Pre-training Music Classification Models via Music Source Separation
by: Garoufis, Christos, et al.
Published: (2023) -
Classical Guitar Duet Separation using GuitarDuets -- a Dataset of Real and Synthesized Guitar Recordings
by: Glytsos, Marios, et al.
Published: (2025) -
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
by: Wang, Ju-Chiang, et al.
Published: (2024) -
voc2vec: A Foundation Model for Non-Verbal Vocalization
by: Koudounas, Alkis, et al.
Published: (2025) -
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
by: Li, Jialu, et al.
Published: (2024)