NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Yen-Ting, Chen, Zhehuai, Zelasko, Piotr, Wan, Zhen, Yang, Xuesong, Chen, Zih-Ching, Puvvada, Krishna C, Fu, Szu-Wei, Hu, Ke, Chiu, Jun Wei, Balam, Jagadeesh, Ginsburg, Boris, Wang, Yu-Chiang Frank, Yang, Chao-Han Huck
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915645943185408
author Lin, Yen-Ting
Chen, Zhehuai
Zelasko, Piotr
Wan, Zhen
Yang, Xuesong
Chen, Zih-Ching
Puvvada, Krishna C
Fu, Szu-Wei
Hu, Ke
Chiu, Jun Wei
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Yang, Chao-Han Huck
author_facet Lin, Yen-Ting
Chen, Zhehuai
Zelasko, Piotr
Wan, Zhen
Yang, Xuesong
Chen, Zih-Ching
Puvvada, Krishna C
Fu, Szu-Wei
Hu, Ke
Chiu, Jun Wei
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Yang, Chao-Han Huck
contents Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The answer would lie in learning dataset-specific features and digesting their knowledge in a single model. Previous methods achieve this by having separate correction language models, resulting in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose a Multi-Task Correction MoE, where we train the experts to become an ``expert'' of speech-to-text, language-to-text and vision-to-text datasets by learning to route each dataset's tokens to its mapped expert. Experiments on the Open ASR Leaderboard show that we explore a new state-of-the-art performance by achieving an average relative 5.0% WER reduction and substantial improvements in BLEU scores for speech and translation tasks. On zero-shot evaluation, NeKo outperforms GPT-3.5 and Claude-Opus with 15.5% to 27.6% relative WER reduction in the Hyporadise benchmark. NeKo performs competitively on grammar and post-OCR correction as a multi-task model.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05945
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
Lin, Yen-Ting
Chen, Zhehuai
Zelasko, Piotr
Wan, Zhen
Yang, Xuesong
Chen, Zih-Ching
Puvvada, Krishna C
Fu, Szu-Wei
Hu, Ke
Chiu, Jun Wei
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Yang, Chao-Han Huck
Computation and Language
Artificial Intelligence
Machine Learning
Multiagent Systems
Audio and Speech Processing
Construction of a general-purpose post-recognition error corrector poses a crucial question: how can we most effectively train a model on a large mixture of domain datasets? The answer would lie in learning dataset-specific features and digesting their knowledge in a single model. Previous methods achieve this by having separate correction language models, resulting in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose a Multi-Task Correction MoE, where we train the experts to become an ``expert'' of speech-to-text, language-to-text and vision-to-text datasets by learning to route each dataset's tokens to its mapped expert. Experiments on the Open ASR Leaderboard show that we explore a new state-of-the-art performance by achieving an average relative 5.0% WER reduction and substantial improvements in BLEU scores for speech and translation tasks. On zero-shot evaluation, NeKo outperforms GPT-3.5 and Claude-Opus with 15.5% to 27.6% relative WER reduction in the Hyporadise benchmark. NeKo performs competitively on grammar and post-OCR correction as a multi-task model.
title NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
topic Computation and Language
Artificial Intelligence
Machine Learning
Multiagent Systems
Audio and Speech Processing
url https://arxiv.org/abs/2411.05945