NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915645943185408 |
|---|---|
| author | Lin, Yen-Ting Chen, Zhehuai Zelasko, Piotr Wan, Zhen Yang, Xuesong Chen, Zih-Ching Puvvada, Krishna C Fu, Szu-Wei Hu, Ke Chiu, Jun Wei Balam, Jagadeesh Ginsburg, Boris Wang, Yu-Chiang Frank Yang, Chao-Han Huck |
| author_facet | Lin, Yen-Ting Chen, Zhehuai Zelasko, Piotr Wan, Zhen Yang, Xuesong Chen, Zih-Ching Puvvada, Krishna C Fu, Szu-Wei Hu, Ke Chiu, Jun Wei Balam, Jagadeesh Ginsburg, Boris Wang, Yu-Chiang Frank Yang, Chao-Han Huck |
| contents | Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The answer would lie in learning dataset-specific features and digesting their knowledge in a single model. Previous methods achieve this by having separate correction language models, resulting in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose a Multi-Task Correction MoE, where we train the experts to become an ``expert'' of speech-to-text, language-to-text and vision-to-text datasets by learning to route each dataset's tokens to its mapped expert. Experiments on the Open ASR Leaderboard show that we explore a new state-of-the-art performance by achieving an average relative 5.0% WER reduction and substantial improvements in BLEU scores for speech and translation tasks. On zero-shot evaluation, NeKo outperforms GPT-3.5 and Claude-Opus with 15.5% to 27.6% relative WER reduction in the Hyporadise benchmark. NeKo performs competitively on grammar and post-OCR correction as a multi-task model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_05945 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model Lin, Yen-Ting Chen, Zhehuai Zelasko, Piotr Wan, Zhen Yang, Xuesong Chen, Zih-Ching Puvvada, Krishna C Fu, Szu-Wei Hu, Ke Chiu, Jun Wei Balam, Jagadeesh Ginsburg, Boris Wang, Yu-Chiang Frank Yang, Chao-Han Huck Computation and Language Artificial Intelligence Machine Learning Multiagent Systems Audio and Speech Processing Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The answer would lie in learning dataset-specific features and digesting their knowledge in a single model. Previous methods achieve this by having separate correction language models, resulting in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose a Multi-Task Correction MoE, where we train the experts to become an ``expert'' of speech-to-text, language-to-text and vision-to-text datasets by learning to route each dataset's tokens to its mapped expert. Experiments on the Open ASR Leaderboard show that we explore a new state-of-the-art performance by achieving an average relative 5.0% WER reduction and substantial improvements in BLEU scores for speech and translation tasks. On zero-shot evaluation, NeKo outperforms GPT-3.5 and Claude-Opus with 15.5% to 27.6% relative WER reduction in the Hyporadise benchmark. NeKo performs competitively on grammar and post-OCR correction as a multi-task model. |
| title | NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model |
| topic | Computation and Language Artificial Intelligence Machine Learning Multiagent Systems Audio and Speech Processing |
| url | https://arxiv.org/abs/2411.05945 |