2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Costa-jussà, Marta R., Yu, Bokai, Andrews, Pierre, Alastruey, Belen, Camgoz, Necati Cihan, Chuang, Joe, Maillard, Jean, Ropers, Christophe, Turkantenko, Arina, Wood, Carleigh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912167090978816
author Costa-jussà, Marta R.
Yu, Bokai
Andrews, Pierre
Alastruey, Belen
Camgoz, Necati Cihan
Chuang, Joe
Maillard, Jean
Ropers, Christophe
Turkantenko, Arina
Wood, Carleigh
author_facet Costa-jussà, Marta R.
Yu, Bokai
Andrews, Pierre
Alastruey, Belen
Camgoz, Necati Cihan
Chuang, Joe
Maillard, Jean
Ropers, Christophe
Turkantenko, Arina
Wood, Carleigh
contents We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the intersection of BELEBELE and FLEURS, and one sign language (ASL). We evaluate 2M-BELEBELE dataset for both 5-shot and zero-shot settings and across languages, the speech comprehension accuracy is ~ 2-3% average lower compared to reading comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
Costa-jussà, Marta R.
Yu, Bokai
Andrews, Pierre
Alastruey, Belen
Camgoz, Necati Cihan
Chuang, Joe
Maillard, Jean
Ropers, Christophe
Turkantenko, Arina
Wood, Carleigh
Computation and Language
I.2.7
We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the intersection of BELEBELE and FLEURS, and one sign language (ASL). We evaluate 2M-BELEBELE dataset for both 5-shot and zero-shot settings and across languages, the speech comprehension accuracy is ~ 2-3% average lower compared to reading comprehension.
title 2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2412.08274