CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Junda, Li, Warren, Novack, Zachary, Namburi, Amit, Chen, Carol, McAuley, Julian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
por: Wu, Junda, et al.
Publicado: (2024)
por: Wu, Junda, et al.
Publicado: (2024)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
por: Mundada, Gagan, et al.
Publicado: (2025)
por: Mundada, Gagan, et al.
Publicado: (2025)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
por: Kim, Haven, et al.
Publicado: (2025)
por: Kim, Haven, et al.
Publicado: (2025)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
por: Long, Phillip, et al.
Publicado: (2024)
por: Long, Phillip, et al.
Publicado: (2024)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
por: Novack, Zachary, et al.
Publicado: (2024)
por: Novack, Zachary, et al.
Publicado: (2024)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
por: Wang, Boyang, et al.
Publicado: (2025)
por: Wang, Boyang, et al.
Publicado: (2025)
Steering Autoregressive Music Generation with Recursive Feature Machines
por: Zhao, Daniel, et al.
Publicado: (2025)
por: Zhao, Daniel, et al.
Publicado: (2025)
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
por: Huang, Jingyue, et al.
Publicado: (2025)
por: Huang, Jingyue, et al.
Publicado: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
por: Novack, Zachary, et al.
Publicado: (2024)
por: Novack, Zachary, et al.
Publicado: (2024)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
por: Kim, Haven, et al.
Publicado: (2026)
por: Kim, Haven, et al.
Publicado: (2026)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
por: Xu, Weihan, et al.
Publicado: (2024)
por: Xu, Weihan, et al.
Publicado: (2024)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
por: Wu, Yusong, et al.
Publicado: (2022)
por: Wu, Yusong, et al.
Publicado: (2022)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
por: Sun, Haoqin, et al.
Publicado: (2025)
por: Sun, Haoqin, et al.
Publicado: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
por: Guinot, Julien, et al.
Publicado: (2025)
por: Guinot, Julien, et al.
Publicado: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
por: Novack, Zachary, et al.
Publicado: (2025)
por: Novack, Zachary, et al.
Publicado: (2025)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
por: Long, Phillip, et al.
Publicado: (2026)
por: Long, Phillip, et al.
Publicado: (2026)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
por: Xiao, Feiyang, et al.
Publicado: (2024)
por: Xiao, Feiyang, et al.
Publicado: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
por: Jing, Xin, et al.
Publicado: (2026)
por: Jing, Xin, et al.
Publicado: (2026)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
por: Tal, Or, et al.
Publicado: (2024)
por: Tal, Or, et al.
Publicado: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
por: Yuan, Yi, et al.
Publicado: (2024)
por: Yuan, Yi, et al.
Publicado: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
por: Primus, Paul, et al.
Publicado: (2025)
por: Primus, Paul, et al.
Publicado: (2025)
Aligning Text-to-Music Evaluation with Human Preferences
por: Huang, Yichen, et al.
Publicado: (2025)
por: Huang, Yichen, et al.
Publicado: (2025)
The Arrow of Time in Music -- Revisiting the Temporal Structure of Music with Distinguishability and Unique Orientability as the Anchor Point
por: Xu, Qi
Publicado: (2023)
por: Xu, Qi
Publicado: (2023)
SoundReactor: Frame-level Online Video-to-Audio Generation
por: Saito, Koichi, et al.
Publicado: (2025)
por: Saito, Koichi, et al.
Publicado: (2025)
LoVA: Long-form Video-to-Audio Generation
por: Cheng, Xin, et al.
Publicado: (2024)
por: Cheng, Xin, et al.
Publicado: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
por: Araz, R. Oguz, et al.
Publicado: (2025)
por: Araz, R. Oguz, et al.
Publicado: (2025)
Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
por: Zhang, Yixiao, et al.
Publicado: (2025)
por: Zhang, Yixiao, et al.
Publicado: (2025)
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
por: Zhang, Yixiao
Publicado: (2024)
por: Zhang, Yixiao
Publicado: (2024)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
por: Surana, Rohan, et al.
Publicado: (2025)
por: Surana, Rohan, et al.
Publicado: (2025)
Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
por: Grötschla, Florian, et al.
Publicado: (2024)
por: Grötschla, Florian, et al.
Publicado: (2024)
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
por: Ciranni, Ruben, et al.
Publicado: (2024)
por: Ciranni, Ruben, et al.
Publicado: (2024)
SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing
por: Chen, Jianyi, et al.
Publicado: (2026)
por: Chen, Jianyi, et al.
Publicado: (2026)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
por: Khanjani, Zahra, et al.
Publicado: (2024)
por: Khanjani, Zahra, et al.
Publicado: (2024)
SIGNL: A Label-Efficient Audio Deepfake Detection System via Spectral-Temporal Graph Non-Contrastive Learning
por: Febrinanto, Falih Gozi, et al.
Publicado: (2025)
por: Febrinanto, Falih Gozi, et al.
Publicado: (2025)
Cacophony: An Improved Contrastive Audio-Text Model
por: Zhu, Ge, et al.
Publicado: (2024)
por: Zhu, Ge, et al.
Publicado: (2024)
Domain Adaptation for Contrastive Audio-Language Models
por: Deshmukh, Soham, et al.
Publicado: (2024)
por: Deshmukh, Soham, et al.
Publicado: (2024)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
por: Rouard, Simon, et al.
Publicado: (2024)
por: Rouard, Simon, et al.
Publicado: (2024)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
por: Zhang, Xueyao, et al.
Publicado: (2023)
por: Zhang, Xueyao, et al.
Publicado: (2023)
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
por: Cheng, Ruoxi, et al.
Publicado: (2024)
por: Cheng, Ruoxi, et al.
Publicado: (2024)
Ejemplares similares
-
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
por: Wu, Junda, et al.
Publicado: (2024) -
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
por: Mundada, Gagan, et al.
Publicado: (2025) -
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
por: Kim, Haven, et al.
Publicado: (2025) -
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
por: Long, Phillip, et al.
Publicado: (2024) -
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
por: Novack, Zachary, et al.
Publicado: (2024)