MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zihao, Liu, Haoxuan, Yu, Jiaxing, Zhang, Tao, Liu, Yan, Zhang, Kejun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
por: Wang, Zihao, et al.
Publicado: (2024)
por: Wang, Zihao, et al.
Publicado: (2024)
SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement
por: Wang, Zihao, et al.
Publicado: (2024)
por: Wang, Zihao, et al.
Publicado: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
por: Zheng, Zihao, et al.
Publicado: (2025)
por: Zheng, Zihao, et al.
Publicado: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
por: Xie, Zeyu, et al.
Publicado: (2025)
por: Xie, Zeyu, et al.
Publicado: (2025)
Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
por: Wang, Ziyu, et al.
Publicado: (2024)
por: Wang, Ziyu, et al.
Publicado: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
por: Xie, Zeyu, et al.
Publicado: (2025)
por: Xie, Zeyu, et al.
Publicado: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
FakeSound: Deepfake General Audio Detection
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
por: Liu, Shansong, et al.
Publicado: (2024)
por: Liu, Shansong, et al.
Publicado: (2024)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
por: You, Fuming, et al.
Publicado: (2024)
por: You, Fuming, et al.
Publicado: (2024)
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
por: Li, Ruiqi, et al.
Publicado: (2024)
por: Li, Ruiqi, et al.
Publicado: (2024)
Can Sound Replace Vision in LLaVA With Token Substitution?
por: Vosoughi, Ali, et al.
Publicado: (2025)
por: Vosoughi, Ali, et al.
Publicado: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
por: Cheripally, Sowmya
Publicado: (2024)
por: Cheripally, Sowmya
Publicado: (2024)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
por: Lv, Yishan, et al.
Publicado: (2026)
por: Lv, Yishan, et al.
Publicado: (2026)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
por: Niizumi, Daisuke, et al.
Publicado: (2024)
por: Niizumi, Daisuke, et al.
Publicado: (2024)
ViMo: Generating Motions from Casual Videos
por: Qiu, Liangdong, et al.
Publicado: (2024)
por: Qiu, Liangdong, et al.
Publicado: (2024)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
por: Xie, Zeyu, et al.
Publicado: (2026)
por: Xie, Zeyu, et al.
Publicado: (2026)
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
por: Xie, Zeyu, et al.
Publicado: (2025)
por: Xie, Zeyu, et al.
Publicado: (2025)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
por: Weck, Benno, et al.
Publicado: (2024)
por: Weck, Benno, et al.
Publicado: (2024)
GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment
por: Wang, Jinting, et al.
Publicado: (2025)
por: Wang, Jinting, et al.
Publicado: (2025)
UniMuMo: Unified Text, Music and Motion Generation
por: Yang, Han, et al.
Publicado: (2024)
por: Yang, Han, et al.
Publicado: (2024)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
por: Li, Junjie, et al.
Publicado: (2024)
por: Li, Junjie, et al.
Publicado: (2024)
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
por: Liu, Haoxuan, et al.
Publicado: (2024)
por: Liu, Haoxuan, et al.
Publicado: (2024)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
por: Salganik, Rebecca, et al.
Publicado: (2026)
por: Salganik, Rebecca, et al.
Publicado: (2026)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
por: Ren, Yong, et al.
Publicado: (2025)
por: Ren, Yong, et al.
Publicado: (2025)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
por: Liu, Shengqiang, et al.
Publicado: (2024)
por: Liu, Shengqiang, et al.
Publicado: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
por: Wang, Anna, et al.
Publicado: (2024)
por: Wang, Anna, et al.
Publicado: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
por: Tian, Wenjie, et al.
Publicado: (2025)
por: Tian, Wenjie, et al.
Publicado: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
por: Ren, Yong, et al.
Publicado: (2024)
por: Ren, Yong, et al.
Publicado: (2024)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
por: Huang, Zhiqi, et al.
Publicado: (2024)
por: Huang, Zhiqi, et al.
Publicado: (2024)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
por: Kantarelis, Spyridon, et al.
Publicado: (2024)
por: Kantarelis, Spyridon, et al.
Publicado: (2024)
pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
por: Devaney, Johanna, et al.
Publicado: (2024)
por: Devaney, Johanna, et al.
Publicado: (2024)
ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal Processing Analysis for Sounds, & Strains Emerging Technology
por: Ali, Ratul, et al.
Publicado: (2023)
por: Ali, Ratul, et al.
Publicado: (2023)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
por: Lin, Yueqian, et al.
Publicado: (2025)
por: Lin, Yueqian, et al.
Publicado: (2025)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
por: Guo, Zixin, et al.
Publicado: (2024)
por: Guo, Zixin, et al.
Publicado: (2024)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
por: Zhang, Liqian, et al.
Publicado: (2024)
por: Zhang, Liqian, et al.
Publicado: (2024)
Ejemplares similares
-
MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
por: Wang, Zihao, et al.
Publicado: (2024) -
SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement
por: Wang, Zihao, et al.
Publicado: (2024) -
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
por: Zheng, Zihao, et al.
Publicado: (2025) -
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
por: Zheng, Zihao, et al.
Publicado: (2026) -
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
por: Xie, Zeyu, et al.
Publicado: (2025)