SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Guinot, Julien, Riou, Alain, Quinton, Elio, Fazekas, György |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
by: Guinot, Julien, et al.
Published: (2025)
by: Guinot, Julien, et al.
Published: (2025)
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
by: Guinot, Julien, et al.
Published: (2024)
by: Guinot, Julien, et al.
Published: (2024)
Semi-Supervised Contrastive Learning of Musical Representations
by: Guinot, Julien, et al.
Published: (2024)
by: Guinot, Julien, et al.
Published: (2024)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
by: Weck, Benno, et al.
Published: (2024)
by: Weck, Benno, et al.
Published: (2024)
Robust Lossy Audio Compression Identification
by: Koops, Hendrik Vincent, et al.
Published: (2024)
by: Koops, Hendrik Vincent, et al.
Published: (2024)
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
by: Mei, Xinhao, et al.
Published: (2026)
by: Mei, Xinhao, et al.
Published: (2026)
Singing Voice Synthesis Using Differentiable LPC and Glottal-Flow-Inspired Wavetables
by: Yu, Chin-Yun, et al.
Published: (2023)
by: Yu, Chin-Yun, et al.
Published: (2023)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
by: Yu, Chin-Yun, et al.
Published: (2024)
by: Yu, Chin-Yun, et al.
Published: (2024)
Exploring Transformer-Based Music Overpainting for Jazz Piano Variations
by: Row, Eleanor, et al.
Published: (2024)
by: Row, Eleanor, et al.
Published: (2024)
Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
by: Zhang, Jincheng, et al.
Published: (2025)
by: Zhang, Jincheng, et al.
Published: (2025)
Composer Style-specific Symbolic Music Generation Using Vector Quantized Discrete Diffusion Models
by: Zhang, Jincheng, et al.
Published: (2023)
by: Zhang, Jincheng, et al.
Published: (2023)
Automatic Music Sample Identification with Multi-Track Contrastive Learning
by: Riou, Alain, et al.
Published: (2025)
by: Riou, Alain, et al.
Published: (2025)
Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
by: Hayes, Ben, et al.
Published: (2025)
by: Hayes, Ben, et al.
Published: (2025)
Music2Latent: Consistency Autoencoders for Latent Audio Compression
by: Pasini, Marco, et al.
Published: (2024)
by: Pasini, Marco, et al.
Published: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
by: Yu, Chin-Yun, et al.
Published: (2022)
by: Yu, Chin-Yun, et al.
Published: (2022)
SLAP: Learning Speaker and Health-Related Representations from Natural Language Supervision
by: Ando, Angelika, et al.
Published: (2025)
by: Ando, Angelika, et al.
Published: (2025)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
by: Yu, Chin-Yun, et al.
Published: (2023)
by: Yu, Chin-Yun, et al.
Published: (2023)
Sound Matching an Analogue Levelling Amplifier Using the Newton-Raphson Method
by: Yu, Chin-Yun, et al.
Published: (2025)
by: Yu, Chin-Yun, et al.
Published: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
by: Riou, Alain, et al.
Published: (2023)
by: Riou, Alain, et al.
Published: (2023)
Time-of-arrival Estimation and Phase Unwrapping of Head-related Transfer Functions With Integer Linear Programming
by: Yu, Chin-Yun, et al.
Published: (2024)
by: Yu, Chin-Yun, et al.
Published: (2024)
Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Exploring trends in audio mixes and masters: Insights from a dataset analysis
by: Mourgela, Angeliki, et al.
Published: (2024)
by: Mourgela, Angeliki, et al.
Published: (2024)
Composers' Evaluations of an AI Music Tool: Insights for Human-Centred Design
by: Row, Eleanor, et al.
Published: (2024)
by: Row, Eleanor, et al.
Published: (2024)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
by: Xiao, Feiyang, et al.
Published: (2024)
by: Xiao, Feiyang, et al.
Published: (2024)
Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
by: Tang, Jingjing, et al.
Published: (2025)
by: Tang, Jingjing, et al.
Published: (2025)
Differentiable All-pole Filters for Time-varying Audio Systems
by: Yu, Chin-Yun, et al.
Published: (2024)
by: Yu, Chin-Yun, et al.
Published: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
by: Jing, Xin, et al.
Published: (2026)
by: Jing, Xin, et al.
Published: (2026)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
by: Wu, Yusong, et al.
Published: (2022)
by: Wu, Yusong, et al.
Published: (2022)
CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
by: Pasini, Marco, et al.
Published: (2025)
by: Pasini, Marco, et al.
Published: (2025)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
by: Araz, R. Oguz, et al.
Published: (2025)
by: Araz, R. Oguz, et al.
Published: (2025)
Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
by: Zhang, Yixiao
Published: (2024)
by: Zhang, Yixiao
Published: (2024)
Can Large Language Models Understand Spatial Audio?
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
by: Sun, Haoqin, et al.
Published: (2025)
by: Sun, Haoqin, et al.
Published: (2025)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
by: Salganik, Rebecca, et al.
Published: (2026)
by: Salganik, Rebecca, et al.
Published: (2026)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
by: Wang, Ziqian, et al.
Published: (2025)
by: Wang, Ziqian, et al.
Published: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
by: Xue, Jinlong, et al.
Published: (2024)
by: Xue, Jinlong, et al.
Published: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
by: Primus, Paul, et al.
Published: (2025)
by: Primus, Paul, et al.
Published: (2025)
Similar Items
-
GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
by: Guinot, Julien, et al.
Published: (2025) -
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
by: Guinot, Julien, et al.
Published: (2024) -
Semi-Supervised Contrastive Learning of Musical Representations
by: Guinot, Julien, et al.
Published: (2024) -
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
by: Weck, Benno, et al.
Published: (2024) -
Robust Lossy Audio Compression Identification
by: Koops, Hendrik Vincent, et al.
Published: (2024)