CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Junda, Li, Warren, Novack, Zachary, Namburi, Amit, Chen, Carol, McAuley, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024)
von: Long, Phillip, et al.
Veröffentlicht: (2024)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
Steering Autoregressive Music Generation with Recursive Feature Machines
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
von: Zhao, Daniel, et al.
Veröffentlicht: (2025)
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
von: Huang, Jingyue, et al.
Veröffentlicht: (2025)
von: Huang, Jingyue, et al.
Veröffentlicht: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
von: Novack, Zachary, et al.
Veröffentlicht: (2025)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
von: Long, Phillip, et al.
Veröffentlicht: (2026)
von: Long, Phillip, et al.
Veröffentlicht: (2026)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
von: Jing, Xin, et al.
Veröffentlicht: (2026)
von: Jing, Xin, et al.
Veröffentlicht: (2026)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
Aligning Text-to-Music Evaluation with Human Preferences
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
The Arrow of Time in Music -- Revisiting the Temporal Structure of Music with Distinguishability and Unique Orientability as the Anchor Point
von: Xu, Qi
Veröffentlicht: (2023)
von: Xu, Qi
Veröffentlicht: (2023)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
LoVA: Long-form Video-to-Audio Generation
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
von: Zhang, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2025)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
von: Zhang, Yixiao
Veröffentlicht: (2024)
von: Zhang, Yixiao
Veröffentlicht: (2024)
Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing
von: Chen, Jianyi, et al.
Veröffentlicht: (2026)
von: Chen, Jianyi, et al.
Veröffentlicht: (2026)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
SIGNL: A Label-Efficient Audio Deepfake Detection System via Spectral-Temporal Graph Non-Contrastive Learning
von: Febrinanto, Falih Gozi, et al.
Veröffentlicht: (2025)
von: Febrinanto, Falih Gozi, et al.
Veröffentlicht: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Domain Adaptation for Contrastive Audio-Language Models
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
von: Wu, Junda, et al.
Veröffentlicht: (2024) -
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2025) -
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025) -
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024) -
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)