HeartMuLa: A Family of Open Sourced Music Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Dongchao, Xie, Yuxin, Yin, Yuguo, Wang, Zheyu, Yi, Xiaoyu, Zhu, Gongxi, Weng, Xiaolong, Xiong, Zihan, Ma, Yingzhe, Cong, Dading, Liu, Jingliang, Huang, Zihang, Ru, Jinghan, Huang, Rongjie, Wan, Haoran, Wang, Peixu, Yu, Kuoxi, Wang, Helin, Liang, Liming, Zhuang, Xianwei, Wang, Yuanyuan, Dingdong, Wang, Guo, Haohan, Cao, Junjie, Ju, Zeqian, Liu, Songxiang, Cao, Yuewen, Weng, Heming, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
von: Yin, Yuguo, et al.
Veröffentlicht: (2025)
von: Yin, Yuguo, et al.
Veröffentlicht: (2025)
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
Artificial Geographically Weighted Neural Network: A Novel Framework for Spatial Analysis with Geographically Weighted Layers
von: Cao, Jianfei, et al.
Veröffentlicht: (2025)
von: Cao, Jianfei, et al.
Veröffentlicht: (2025)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Do we really have to filter out random noise in pre-training data for language models?
von: Ru, Jinghan, et al.
Veröffentlicht: (2025)
von: Ru, Jinghan, et al.
Veröffentlicht: (2025)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
von: Ru, Jinghan, et al.
Veröffentlicht: (2026)
von: Ru, Jinghan, et al.
Veröffentlicht: (2026)
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
von: Ma, Yingzhe, et al.
Veröffentlicht: (2026)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
von: Fu, Siyuan, et al.
Veröffentlicht: (2025)
von: Fu, Siyuan, et al.
Veröffentlicht: (2025)
SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
von: Luo, Jiehui, et al.
Veröffentlicht: (2025)
von: Luo, Jiehui, et al.
Veröffentlicht: (2025)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
von: You, Fuming, et al.
Veröffentlicht: (2024)
von: You, Fuming, et al.
Veröffentlicht: (2024)
MIGA: Mixture-of-Experts with Group Aggregation for Stock Market Prediction
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
Composite phenotypes and multiple organ systems aging clocks predict cognitive decline
von: Yingzhe Wang
Veröffentlicht: (2025)
von: Yingzhe Wang
Veröffentlicht: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
von: Yang, Jinghan, et al.
Veröffentlicht: (2025)
von: Yang, Jinghan, et al.
Veröffentlicht: (2025)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Unveiling the dynamic evolution of innovation networks in emerging economies: A complex network approach
von: Zeqian Wang, et al.
Veröffentlicht: (2024)
von: Zeqian Wang, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
von: Cao, Xianwei, et al.
Veröffentlicht: (2026)
von: Cao, Xianwei, et al.
Veröffentlicht: (2026)
Default Contagion, Matrix Approximation, and Control in Sparse Financial Networks
von: Zhang, Aoxin, et al.
Veröffentlicht: (2026)
von: Zhang, Aoxin, et al.
Veröffentlicht: (2026)
Fortunellin ameliorates LPS‐induced acute lung injury, inflammation, and collagen deposition by restraining the TLR4/NF‐κB/NLRP3 pathway
von: Danjuan Liu, et al.
Veröffentlicht: (2024)
von: Danjuan Liu, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning-aided Transmission Design for Energy-efficient Link Optimization in Vehicular Communications
von: Wang, Zhengpeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhengpeng, et al.
Veröffentlicht: (2024)
Homogeneous fractional integral operators on weighted Lebesgue, Morrey and Campanato spaces
von: Du, Jingliang, et al.
Veröffentlicht: (2025)
von: Du, Jingliang, et al.
Veröffentlicht: (2025)
Influence of Interstitial Segregation on Grain Boundary Cohesion in Ferritic Steels
von: Jingliang Wang, et al.
Veröffentlicht: (2026)
von: Jingliang Wang, et al.
Veröffentlicht: (2026)
Bound states and atomic interaction in giant atom waveguide QED with dispersive coupling
von: Weng, Mingzhu, et al.
Veröffentlicht: (2024)
von: Weng, Mingzhu, et al.
Veröffentlicht: (2024)
A mean curvature type flow with capillary boundary in a unit ball
von: Wang, Guofang, et al.
Veröffentlicht: (2020)
von: Wang, Guofang, et al.
Veröffentlicht: (2020)
Beamforming-Codebook-Aware Channel Knowledge Map Construction for Multi-Antenna Systems
von: Wang, Haohan, et al.
Veröffentlicht: (2025)
von: Wang, Haohan, et al.
Veröffentlicht: (2025)
PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance
von: Weng, Haohan, et al.
Veröffentlicht: (2024)
von: Weng, Haohan, et al.
Veröffentlicht: (2024)
BeamCKM: A Framework of Channel Knowledge Map Construction for Multi-Antenna Systems
von: Wang, Haohan, et al.
Veröffentlicht: (2025)
von: Wang, Haohan, et al.
Veröffentlicht: (2025)
DistDD: Distributed Data Distillation Aggregation through Gradient Matching
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
von: Wang, Peiran, et al.
Veröffentlicht: (2024)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
Interaction and entanglement engineering in driven giant atoms setup with coupled resonator waveguide
von: Weng, Mingzhu, et al.
Veröffentlicht: (2024)
von: Weng, Mingzhu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024) -
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
von: Yang, Dongchao, et al.
Veröffentlicht: (2026) -
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025) -
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
von: Yin, Yuguo, et al.
Veröffentlicht: (2025) -
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)