_version_ 1866914208755482624
author An, Keyu
Chen, Yanni
Chen, Zhigao
Deng, Chong
Du, Zhihao
Gao, Changfeng
Gao, Zhifu
Gong, Bo
Li, Xiangang
Li, Yabin
Liu, Ying
Lv, Xiang
Ji, Yunjie
Jiang, Yiheng
Ma, Bin
Luo, Haoneng
Ni, Chongjia
Pan, Zexu
Peng, Yiping
Peng, Zhendong
Wang, Peiyao
Wang, Hao
Wang, Haoxu
Wang, Wen
Wang, Wupeng
Wu, Yuzhong
Tian, Biao
Tan, Zhentao
Yang, Nan
Yuan, Bin
Ye, Jieping
Yu, Jixing
Zhang, Qinglin
Zou, Kun
Zhao, Han
Zhao, Shengkui
Zhou, Jingren
Zhu, Yanqiao
author_facet An, Keyu
Chen, Yanni
Chen, Zhigao
Deng, Chong
Du, Zhihao
Gao, Changfeng
Gao, Zhifu
Gong, Bo
Li, Xiangang
Li, Yabin
Liu, Ying
Lv, Xiang
Ji, Yunjie
Jiang, Yiheng
Ma, Bin
Luo, Haoneng
Ni, Chongjia
Pan, Zexu
Peng, Yiping
Peng, Zhendong
Wang, Peiyao
Wang, Hao
Wang, Haoxu
Wang, Wen
Wang, Wupeng
Wu, Yuzhong
Tian, Biao
Tan, Zhentao
Yang, Nan
Yuan, Bin
Ye, Jieping
Yu, Jixing
Zhang, Qinglin
Zou, Kun
Zhao, Han
Zhao, Shengkui
Zhou, Jingren
Zhu, Yanqiao
contents In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR .
format Preprint
id arxiv_https___arxiv_org_abs_2509_12508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fun-ASR Technical Report
An, Keyu
Chen, Yanni
Chen, Zhigao
Deng, Chong
Du, Zhihao
Gao, Changfeng
Gao, Zhifu
Gong, Bo
Li, Xiangang
Li, Yabin
Liu, Ying
Lv, Xiang
Ji, Yunjie
Jiang, Yiheng
Ma, Bin
Luo, Haoneng
Ni, Chongjia
Pan, Zexu
Peng, Yiping
Peng, Zhendong
Wang, Peiyao
Wang, Hao
Wang, Haoxu
Wang, Wen
Wang, Wupeng
Wu, Yuzhong
Tian, Biao
Tan, Zhentao
Yang, Nan
Yuan, Bin
Ye, Jieping
Yu, Jixing
Zhang, Qinglin
Zou, Kun
Zhao, Han
Zhao, Shengkui
Zhou, Jingren
Zhu, Yanqiao
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR .
title Fun-ASR Technical Report
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.12508