MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huzaifah, Muhammad, Lin, Geyu, Liu, Tianchi, Sailor, Hardik B., Tan, Kye Min, Vangani, Tarun K., Wang, Qiongqiong, Wong, Jeremy H. M., Wu, Jinyang, Chen, Nancy F., Aw, Ai Ti
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912580927225856
author Huzaifah, Muhammad
Lin, Geyu
Liu, Tianchi
Sailor, Hardik B.
Tan, Kye Min
Vangani, Tarun K.
Wang, Qiongqiong
Wong, Jeremy H. M.
Wu, Jinyang
Chen, Nancy F.
Aw, Ai Ti
author_facet Huzaifah, Muhammad
Lin, Geyu
Liu, Tianchi
Sailor, Hardik B.
Tan, Kye Min
Vangani, Tarun K.
Wang, Qiongqiong
Wong, Jeremy H. M.
Wu, Jinyang
Chen, Nancy F.
Aw, Ai Ti
contents This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications. Developed as part of Singapore's National Multimodal Large Language Model Programme, the MERaLiON-SpeechEncoder is tailored to address the speech processing needs in Singapore and the surrounding Southeast Asian region. The model currently supports mainly English, including the variety spoken in Singapore. We are actively expanding our datasets to gradually cover other languages in subsequent releases. The MERaLiON-SpeechEncoder was pre-trained from scratch on 200,000 hours of unlabelled speech data using a self-supervised learning approach based on masked language modelling. We describe our training procedure and hyperparameter tuning experiments in detail below. Our evaluation demonstrates improvements to spontaneous and Singapore speech benchmarks for speech recognition, while remaining competitive to other state-of-the-art speech encoders across ten other speech tasks. We commit to releasing our model, supporting broader research endeavours, both in Singapore and beyond.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
Huzaifah, Muhammad
Lin, Geyu
Liu, Tianchi
Sailor, Hardik B.
Tan, Kye Min
Vangani, Tarun K.
Wang, Qiongqiong
Wong, Jeremy H. M.
Wu, Jinyang
Chen, Nancy F.
Aw, Ai Ti
Computation and Language
Artificial Intelligence
Audio and Speech Processing
This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications. Developed as part of Singapore's National Multimodal Large Language Model Programme, the MERaLiON-SpeechEncoder is tailored to address the speech processing needs in Singapore and the surrounding Southeast Asian region. The model currently supports mainly English, including the variety spoken in Singapore. We are actively expanding our datasets to gradually cover other languages in subsequent releases. The MERaLiON-SpeechEncoder was pre-trained from scratch on 200,000 hours of unlabelled speech data using a self-supervised learning approach based on masked language modelling. We describe our training procedure and hyperparameter tuning experiments in detail below. Our evaluation demonstrates improvements to spontaneous and Singapore speech benchmarks for speech recognition, while remaining competitive to other state-of-the-art speech encoders across ten other speech tasks. We commit to releasing our model, supporting broader research endeavours, both in Singapore and beyond.
title MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.11538