Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Wei, Yang, Xiaoyu, Yao, Zengwei, Kuang, Fangjun, Yang, Yifan, Guo, Liyong, Lin, Long, Povey, Daniel
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929209383845888
author Kang, Wei
Yang, Xiaoyu
Yao, Zengwei
Kuang, Fangjun
Yang, Yifan
Guo, Liyong
Lin, Long
Povey, Daniel
author_facet Kang, Wei
Yang, Xiaoyu
Yao, Zengwei
Kuang, Fangjun
Yang, Yifan
Guo, Liyong
Lin, Long
Povey, Daniel
contents In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox. To the best of our knowledge, Libriheavy is the largest freely-available corpus of speech with supervisions. Different from other open-sourced datasets that only provide normalized transcriptions, Libriheavy contains richer information such as punctuation, casing and text context, which brings more flexibility for system building. Specifically, we propose a general and efficient pipeline to locate, align and segment the audios in previously published Librilight to its corresponding texts. The same as Librilight, Libriheavy also has three training subsets small, medium, large of the sizes 500h, 5000h, 50000h respectively. We also extract the dev and test evaluation sets from the aligned audios and guarantee there is no overlapping speakers and books in training sets. Baseline systems are built on the popular CTC-Attention and transducer models. Additionally, we open-source our dataset creatation pipeline which can also be used to other audio alignment tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08105
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
Kang, Wei
Yang, Xiaoyu
Yao, Zengwei
Kuang, Fangjun
Yang, Yifan
Guo, Liyong
Lin, Long
Povey, Daniel
Audio and Speech Processing
Sound
In this paper, we introduce Libriheavy, a large-scale ASR corpus consisting of 50,000 hours of read English speech derived from LibriVox. To the best of our knowledge, Libriheavy is the largest freely-available corpus of speech with supervisions. Different from other open-sourced datasets that only provide normalized transcriptions, Libriheavy contains richer information such as punctuation, casing and text context, which brings more flexibility for system building. Specifically, we propose a general and efficient pipeline to locate, align and segment the audios in previously published Librilight to its corresponding texts. The same as Librilight, Libriheavy also has three training subsets small, medium, large of the sizes 500h, 5000h, 50000h respectively. We also extract the dev and test evaluation sets from the aligned audios and guarantee there is no overlapping speakers and books in training sets. Baseline systems are built on the popular CTC-Attention and transducer models. Additionally, we open-source our dataset creatation pipeline which can also be used to other audio alignment tasks.
title Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.08105