Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kashiwagi, Yosuke, Futami, Hayato, Tsunoo, Emiru, Asakawa, Satoshi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912409398018048
author Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Asakawa, Satoshi
author_facet Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Asakawa, Satoshi
contents This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
Kashiwagi, Yosuke
Futami, Hayato
Tsunoo, Emiru
Asakawa, Satoshi
Computation and Language
Audio and Speech Processing
This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1.
title Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.01439