Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jiafeng, Dong, Yuanliang, Liu, Hongjia, Cheng, Yuqing, Guo, Zhancheng, Liang, Huijing, Zhan, Wenbo, Sun, Yuming, Li, Xiaobing, Yu, Feng, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
von: Wu, Shangda, et al.
Veröffentlicht: (2025)
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
von: Wu, Shangda, et al.
Veröffentlicht: (2024)
Exploring Tokenization Methods for Multitrack Sheet Music Generation
von: Wang, Yashan, et al.
Veröffentlicht: (2024)
von: Wang, Yashan, et al.
Veröffentlicht: (2024)
Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation
von: Cheng, Yuqing, et al.
Veröffentlicht: (2026)
von: Cheng, Yuqing, et al.
Veröffentlicht: (2026)
NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms
von: Wang, Yashan, et al.
Veröffentlicht: (2025)
von: Wang, Yashan, et al.
Veröffentlicht: (2025)
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
von: Dai, Congren, et al.
Veröffentlicht: (2025)
von: Dai, Congren, et al.
Veröffentlicht: (2025)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
von: Liu, Weide, et al.
Veröffentlicht: (2023)
von: Liu, Weide, et al.
Veröffentlicht: (2023)
ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification
von: Sun, Bochao, et al.
Veröffentlicht: (2025)
von: Sun, Bochao, et al.
Veröffentlicht: (2025)
TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation
von: Feng, Yongsheng, et al.
Veröffentlicht: (2025)
von: Feng, Yongsheng, et al.
Veröffentlicht: (2025)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
von: Lam, Max W. Y., et al.
Veröffentlicht: (2025)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
von: Liu, Cheng, et al.
Veröffentlicht: (2025)
Acoustic Overspecification in Electronic Dance Music Taxonomy
von: Xu, Weilun, et al.
Veröffentlicht: (2025)
von: Xu, Weilun, et al.
Veröffentlicht: (2025)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
von: Zhang, Dengming, et al.
Veröffentlicht: (2024)
von: Zhang, Dengming, et al.
Veröffentlicht: (2024)
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
von: Li, Haowen, et al.
Veröffentlicht: (2026)
von: Li, Haowen, et al.
Veröffentlicht: (2026)
Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
von: Wang, Kangdi, et al.
Veröffentlicht: (2025)
von: Wang, Kangdi, et al.
Veröffentlicht: (2025)
High-Fidelity Music Vocoder using Neural Audio Codecs
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
EMelodyGen: Emotion-Conditioned Melody Generation in ABC Notation with the Musical Feature Template
von: Zhou, Monan, et al.
Veröffentlicht: (2023)
von: Zhou, Monan, et al.
Veröffentlicht: (2023)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing
von: Li, Yin, et al.
Veröffentlicht: (2024)
von: Li, Yin, et al.
Veröffentlicht: (2024)
Hankel-FNO: Fast Underwater Acoustic Charting Via Physics-Encoded Fourier Neural Operator
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
The Sound Demixing Challenge 2023 $\unicode{x2013}$ Music Demixing Track
von: Fabbro, Giorgio, et al.
Veröffentlicht: (2023)
von: Fabbro, Giorgio, et al.
Veröffentlicht: (2023)
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
von: Wang, Yanyun, et al.
Veröffentlicht: (2026)
von: Wang, Yanyun, et al.
Veröffentlicht: (2026)
ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation
von: Liu, Yakun, et al.
Veröffentlicht: (2026)
von: Liu, Yakun, et al.
Veröffentlicht: (2026)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation
von: Yoo, HaeJun, et al.
Veröffentlicht: (2024)
von: Yoo, HaeJun, et al.
Veröffentlicht: (2024)
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
von: Wu, Shangda, et al.
Veröffentlicht: (2024) -
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
von: Wu, Shangda, et al.
Veröffentlicht: (2025) -
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
von: Wu, Shangda, et al.
Veröffentlicht: (2024) -
Exploring Tokenization Methods for Multitrack Sheet Music Generation
von: Wang, Yashan, et al.
Veröffentlicht: (2024) -
Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation
von: Cheng, Yuqing, et al.
Veröffentlicht: (2026)