Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912409398018048 |
|---|---|
| author | Kashiwagi, Yosuke Futami, Hayato Tsunoo, Emiru Asakawa, Satoshi |
| author_facet | Kashiwagi, Yosuke Futami, Hayato Tsunoo, Emiru Asakawa, Satoshi |
| contents | This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_01439 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Kashiwagi, Yosuke Futami, Hayato Tsunoo, Emiru Asakawa, Satoshi Computation and Language Audio and Speech Processing This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates w2v-BERT self-supervised model, an encoder-decoder backbone built on E-Branchformer, and a joint CTC-attention decoding strategy. The training corpus comprises varied speech data, of not only public corpora but also in-house data, thereby enhancing the model's robustness to different speaking styles and acoustic conditions. Through evaluations on multiple benchmarks, Whale achieved comparable performance to existing models. In particular, it achieves a word error rate of 2.4% on the Librispeech test-clean set and a character error rate of 3.4% on the CSJ eval3 set, outperforming Whisper large-v3 and OWSM v3.1. |
| title | Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.01439 |