FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Keyu, Chen, Qian, Deng, Chong, Du, Zhihao, Gao, Changfeng, Gao, Zhifu, Gu, Yue, He, Ting, Hu, Hangrui, Hu, Kai, Ji, Shengpeng, Li, Yabin, Li, Zerui, Lu, Heng, Luo, Haoneng, Lv, Xiang, Ma, Bin, Ma, Ziyang, Ni, Chongjia, Song, Changhe, Shi, Jiaqi, Shi, Xian, Wang, Hao, Wang, Wen, Wang, Yuxuan, Xiao, Zhangyu, Yan, Zhijie, Yang, Yexin, Zhang, Bin, Zhang, Qinglin, Zhang, Shiliang, Zhao, Nan, Zheng, Siqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Fun-ASR Technical Report
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
di: Chen, Qian, et al.
Pubblicazione: (2025)
di: Chen, Qian, et al.
Pubblicazione: (2025)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Differentiable Reward Optimization for LLM based TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025)
di: Du, Zhihao, et al.
Pubblicazione: (2025)
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
di: Du, Zhihao, et al.
Pubblicazione: (2023)
di: Du, Zhihao, et al.
Pubblicazione: (2023)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
di: Zhang, Chong, et al.
Pubblicazione: (2025)
di: Zhang, Chong, et al.
Pubblicazione: (2025)
Misclassification Rate and Privacy-Utility Trade-offs in Graph Convolutional Networks via Subsampling Stability
di: Zhang, Yexin, et al.
Pubblicazione: (2026)
di: Zhang, Yexin, et al.
Pubblicazione: (2026)
Graph Attention is Not Always Beneficial: A Theoretical Analysis of Graph Attention Mechanisms via Contextual Stochastic Block Models
di: Ma, Zhongtian, et al.
Pubblicazione: (2024)
di: Ma, Zhongtian, et al.
Pubblicazione: (2024)
A dual‐band dual‐polarized omnidirectional patch antenna for on‐body application
di: Guo‐Ping Gao, et al.
Pubblicazione: (2024)
di: Guo‐Ping Gao, et al.
Pubblicazione: (2024)
On components of the tensor square of a Weyl module
di: Gao, Shiliang, et al.
Pubblicazione: (2023)
di: Gao, Shiliang, et al.
Pubblicazione: (2023)
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
di: Peng, Yizhou, et al.
Pubblicazione: (2025)
di: Peng, Yizhou, et al.
Pubblicazione: (2025)
Deriving the Energy Function of Non-repeaters from CHIME/FRB Baseband Data
di: Ma, Wenqi, et al.
Pubblicazione: (2025)
di: Ma, Wenqi, et al.
Pubblicazione: (2025)
Fun-Audio-Chat Technical Report
di: Tongyi Fun Team, et al.
Pubblicazione: (2025)
di: Tongyi Fun Team, et al.
Pubblicazione: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
di: Yu, Fan, et al.
Pubblicazione: (2024)
di: Yu, Fan, et al.
Pubblicazione: (2024)
Community Detection in the Multi-View Stochastic Block Model
di: Zhang, Yexin, et al.
Pubblicazione: (2024)
di: Zhang, Yexin, et al.
Pubblicazione: (2024)
Gated Graph Attention Networks with Learnable Temperature
di: Ma, Zhongtian, et al.
Pubblicazione: (2026)
di: Ma, Zhongtian, et al.
Pubblicazione: (2026)
Ultra-short-term solar power forecasting by deep learning and data reconstruction
di: Wang, Jinbao, et al.
Pubblicazione: (2025)
di: Wang, Jinbao, et al.
Pubblicazione: (2025)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
di: Zhao, Shengkui, et al.
Pubblicazione: (2023)
di: Zhao, Shengkui, et al.
Pubblicazione: (2023)
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
di: Gao, Bin-Bin, et al.
Pubblicazione: (2026)
di: Gao, Bin-Bin, et al.
Pubblicazione: (2026)
Aggregation‐Induced Electronic Modulation of Carbon Nitride Nanosheets for Broadband Solar Hydrogen Production
di: Xinning Ma, et al.
Pubblicazione: (2026)
di: Xinning Ma, et al.
Pubblicazione: (2026)
Clarifying the clinical implications of clonal hematopoiesis in late‐onset seropositive rheumatoid arthritis : comment on the article by Zhao et al
di: Weiqi Hu, et al.
Pubblicazione: (2026)
di: Weiqi Hu, et al.
Pubblicazione: (2026)
Distributional Causal Mediation via Conditional Generative Modeling
di: Zhang, Jinlun, et al.
Pubblicazione: (2026)
di: Zhang, Jinlun, et al.
Pubblicazione: (2026)
Anisotropic Gauss Reconstruction for Unoriented Point Clouds
di: Ma, Yueji, et al.
Pubblicazione: (2024)
di: Ma, Yueji, et al.
Pubblicazione: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
di: Cheng, Luyao, et al.
Pubblicazione: (2023)
di: Cheng, Luyao, et al.
Pubblicazione: (2023)
Degrees of the stretched Kostka quasi‐polynomials
di: Shiliang Gao, et al.
Pubblicazione: (2024)
di: Shiliang Gao, et al.
Pubblicazione: (2024)
Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
di: Zhang, Jingyao, et al.
Pubblicazione: (2025)
di: Zhang, Jingyao, et al.
Pubblicazione: (2025)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
di: Chen, Qian, et al.
Pubblicazione: (2023)
di: Chen, Qian, et al.
Pubblicazione: (2023)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
di: Chen, Qian, et al.
Pubblicazione: (2024)
di: Chen, Qian, et al.
Pubblicazione: (2024)
LF-PGVIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras using Points and Geodesic Segments
di: Wang, Ze, et al.
Pubblicazione: (2023)
di: Wang, Ze, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024) -
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
di: Du, Zhihao, et al.
Pubblicazione: (2024) -
Fun-ASR Technical Report
di: An, Keyu, et al.
Pubblicazione: (2025) -
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025) -
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)