Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Manco, Ilaria, Salamon, Justin, Nieto, Oriol |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Text2midi: Generating Symbolic Music from Captions
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024)
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
von: Changin, Choi, et al.
Veröffentlicht: (2024)
von: Changin, Choi, et al.
Veröffentlicht: (2024)
Classifier-Guided Captioning Across Modalities
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
von: Shaulov, Ariel, et al.
Veröffentlicht: (2025)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
von: Prajwal, K R, et al.
Veröffentlicht: (2024)
von: Prajwal, K R, et al.
Veröffentlicht: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
The Interpretation Gap in Text-to-Music Generation Models
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Aligning Text-to-Music Evaluation with Human Preferences
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
von: Tsai, Fang-Duo, et al.
Veröffentlicht: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
Cross-Modal Learning for Music-to-Music-Video Description Generation
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
A Human-in-the-Loop Approach to Improving Cross-Text Prosody Transfer
von: Maurya, Himanshu, et al.
Veröffentlicht: (2024)
von: Maurya, Himanshu, et al.
Veröffentlicht: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
von: Baskar, Murali Karthick, et al.
Veröffentlicht: (2024)
von: Baskar, Murali Karthick, et al.
Veröffentlicht: (2024)
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Musical ethnocentrism in Large Language Models
von: Kruspe, Anna
Veröffentlicht: (2025)
von: Kruspe, Anna
Veröffentlicht: (2025)
DAIRHuM: A Platform for Directly Aligning AI Representations with Human Musical Judgments applied to Carnatic Music
von: Ravikumar, Prashanth Thattai
Veröffentlicht: (2024)
von: Ravikumar, Prashanth Thattai
Veröffentlicht: (2024)
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
von: Liu, Shuyang, et al.
Veröffentlicht: (2025)
von: Liu, Shuyang, et al.
Veröffentlicht: (2025)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)
von: Luong, Justin, et al.
Veröffentlicht: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
Tuning Music Education: AI-Powered Personalization in Learning Music
von: Sanganeria, Mayank, et al.
Veröffentlicht: (2024)
von: Sanganeria, Mayank, et al.
Veröffentlicht: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
Sing it, Narrate it: Quality Musical Lyrics Translation
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
von: Go, Gyehun, et al.
Veröffentlicht: (2025)
von: Go, Gyehun, et al.
Veröffentlicht: (2025)
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024) -
Text2midi: Generating Symbolic Music from Captions
von: Bhandari, Keshav, et al.
Veröffentlicht: (2024) -
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023) -
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023) -
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)