InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chong, Ma, Yukun, Chen, Qian, Wang, Wen, Zhao, Shengkui, Pan, Zexu, Wang, Hao, Ni, Chongjia, Nguyen, Trung Hieu, Zhou, Kun, Jiang, Yidi, Tan, Chaohong, Gao, Zhifu, Du, Zhihao, Ma, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
SPGM: Prioritizing Local Features for enhanced speech separation performance
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
von: Pan, Zexu, et al.
Veröffentlicht: (2026)
von: Pan, Zexu, et al.
Veröffentlicht: (2026)
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2026)
von: Pan, Zexu, et al.
Veröffentlicht: (2026)
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
Phase Repair for Time-Domain Convolutional Neural Networks in Music Super-Resolution
von: Zhang, Yenan, et al.
Veröffentlicht: (2023)
von: Zhang, Yenan, et al.
Veröffentlicht: (2023)
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
von: Ni-Hahn, Stephen, et al.
Veröffentlicht: (2025)
von: Ni-Hahn, Stephen, et al.
Veröffentlicht: (2025)
Evaluating High-Resolution Piano Sustain Pedal Depth Estimation with Musically Informed Metrics
von: Zhang, Hanwen, et al.
Veröffentlicht: (2025)
von: Zhang, Hanwen, et al.
Veröffentlicht: (2025)
Assessing Data Replication in Symbolic Music via Adapted Structural Similarity Index Measure
von: Ji, Shulei, et al.
Veröffentlicht: (2025)
von: Ji, Shulei, et al.
Veröffentlicht: (2025)
Large Language Models: From Notes to Musical Form
von: Atassi, Lilac
Veröffentlicht: (2024)
von: Atassi, Lilac
Veröffentlicht: (2024)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
Music2Fail: Transfer Music to Failed Recorder Style
von: Leong, Chon In, et al.
Veröffentlicht: (2024)
von: Leong, Chon In, et al.
Veröffentlicht: (2024)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
Pre-training Music Classification Models via Music Source Separation
von: Garoufis, Christos, et al.
Veröffentlicht: (2023)
von: Garoufis, Christos, et al.
Veröffentlicht: (2023)
Musical Word Embedding for Music Tagging and Retrieval
von: Doh, SeungHeon, et al.
Veröffentlicht: (2024)
von: Doh, SeungHeon, et al.
Veröffentlicht: (2024)
Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
FakeMusicCaps: a Dataset for Detection and Attribution of Synthetic Music Generated via Text-to-Music Models
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
von: Fang, Kun, et al.
Veröffentlicht: (2025)
von: Fang, Kun, et al.
Veröffentlicht: (2025)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
MusicMamba: A Dual-Feature Modeling Approach for Generating Chinese Traditional Music with Modal Precision
von: Chen, Jiatao, et al.
Veröffentlicht: (2024)
von: Chen, Jiatao, et al.
Veröffentlicht: (2024)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Music Style Transfer with Time-Varying Inversion of Diffusion Models
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
Lead Instrument Detection from Multitrack Music
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
von: Ou, Longshen, et al.
Veröffentlicht: (2025)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
YuE: Scaling Open Foundation Models for Long-Form Music Generation
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
von: Yuan, Ruibin, et al.
Veröffentlicht: (2025)
The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
von: Li, Jiajia, et al.
Veröffentlicht: (2024)
von: Li, Jiajia, et al.
Veröffentlicht: (2024)
Expressive Timing in Hindustani Vocal Music
von: Bhake, Yash, et al.
Veröffentlicht: (2025)
von: Bhake, Yash, et al.
Veröffentlicht: (2025)
AILive Mixer: A Deep Learning based Zero Latency Automatic Music Mixer for Live Music Performances
von: Zurale, Devansh, et al.
Veröffentlicht: (2026)
von: Zurale, Devansh, et al.
Veröffentlicht: (2026)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
von: Hao, Chunbo, et al.
Veröffentlicht: (2025)
von: Hao, Chunbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025) -
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023) -
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024) -
SPGM: Prioritizing Local Features for enhanced speech separation performance
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023) -
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)