Fun-ASR Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914208755482624 |
|---|---|
| author | An, Keyu Chen, Yanni Chen, Zhigao Deng, Chong Du, Zhihao Gao, Changfeng Gao, Zhifu Gong, Bo Li, Xiangang Li, Yabin Liu, Ying Lv, Xiang Ji, Yunjie Jiang, Yiheng Ma, Bin Luo, Haoneng Ni, Chongjia Pan, Zexu Peng, Yiping Peng, Zhendong Wang, Peiyao Wang, Hao Wang, Haoxu Wang, Wen Wang, Wupeng Wu, Yuzhong Tian, Biao Tan, Zhentao Yang, Nan Yuan, Bin Ye, Jieping Yu, Jixing Zhang, Qinglin Zou, Kun Zhao, Han Zhao, Shengkui Zhou, Jingren Zhu, Yanqiao |
| author_facet | An, Keyu Chen, Yanni Chen, Zhigao Deng, Chong Du, Zhihao Gao, Changfeng Gao, Zhifu Gong, Bo Li, Xiangang Li, Yabin Liu, Ying Lv, Xiang Ji, Yunjie Jiang, Yiheng Ma, Bin Luo, Haoneng Ni, Chongjia Pan, Zexu Peng, Yiping Peng, Zhendong Wang, Peiyao Wang, Hao Wang, Haoxu Wang, Wen Wang, Wupeng Wu, Yuzhong Tian, Biao Tan, Zhentao Yang, Nan Yuan, Bin Ye, Jieping Yu, Jixing Zhang, Qinglin Zou, Kun Zhao, Han Zhao, Shengkui Zhou, Jingren Zhu, Yanqiao |
| contents | In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_12508 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fun-ASR Technical Report An, Keyu Chen, Yanni Chen, Zhigao Deng, Chong Du, Zhihao Gao, Changfeng Gao, Zhifu Gong, Bo Li, Xiangang Li, Yabin Liu, Ying Lv, Xiang Ji, Yunjie Jiang, Yiheng Ma, Bin Luo, Haoneng Ni, Chongjia Pan, Zexu Peng, Yiping Peng, Zhendong Wang, Peiyao Wang, Hao Wang, Haoxu Wang, Wen Wang, Wupeng Wu, Yuzhong Tian, Biao Tan, Zhentao Yang, Nan Yuan, Bin Ye, Jieping Yu, Jixing Zhang, Qinglin Zou, Kun Zhao, Han Zhao, Shengkui Zhou, Jingren Zhu, Yanqiao Computation and Language Artificial Intelligence Sound Audio and Speech Processing In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR . |
| title | Fun-ASR Technical Report |
| topic | Computation and Language Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.12508 |