Speak & Improve Challenge 2025: Tasks and Baseline Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Mengjie, Knill, Kate, Banno, Stefano, Tang, Siyuan, Karanasou, Penny, Gales, Mark J. F., Nicholls, Diane
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916528207691776
author Qian, Mengjie
Knill, Kate
Banno, Stefano
Tang, Siyuan
Karanasou, Penny
Gales, Mark J. F.
Nicholls, Diane
author_facet Qian, Mengjie
Knill, Kate
Banno, Stefano
Tang, Siyuan
Karanasou, Penny
Gales, Mark J. F.
Nicholls, Diane
contents This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment and feedback, with tasks associated with both the underlying technology and language learning feedback. Linked with the challenge, the Speak & Improve (S&I) Corpus 2025 is being pre-released, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The corpus consists of approximately 315 hours of audio data from second language English learners with holistic scores, and a 55-hour subset with manual transcriptions and error labels. The Challenge has four shared tasks: Automatic Speech Recognition (ASR), Spoken Language Assessment (SLA), Spoken Grammatical Error Correction (SGEC), and Spoken Grammatical Error Correction Feedback (SGECF). Each of these tasks has a closed track where a predetermined set of models and data sources are allowed to be used, and an open track where any public resource may be used. Challenge participants may do one or more of the tasks. This paper describes the challenge, the S&I Corpus 2025, and the baseline systems released for the Challenge.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11985
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speak & Improve Challenge 2025: Tasks and Baseline Systems
Qian, Mengjie
Knill, Kate
Banno, Stefano
Tang, Siyuan
Karanasou, Penny
Gales, Mark J. F.
Nicholls, Diane
Computation and Language
This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment and feedback, with tasks associated with both the underlying technology and language learning feedback. Linked with the challenge, the Speak & Improve (S&I) Corpus 2025 is being pre-released, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The corpus consists of approximately 315 hours of audio data from second language English learners with holistic scores, and a 55-hour subset with manual transcriptions and error labels. The Challenge has four shared tasks: Automatic Speech Recognition (ASR), Spoken Language Assessment (SLA), Spoken Grammatical Error Correction (SGEC), and Spoken Grammatical Error Correction Feedback (SGECF). Each of these tasks has a closed track where a predetermined set of models and data sources are allowed to be used, and an open track where any public resource may be used. Challenge participants may do one or more of the tasks. This paper describes the challenge, the S&I Corpus 2025, and the baseline systems released for the Challenge.
title Speak & Improve Challenge 2025: Tasks and Baseline Systems
topic Computation and Language
url https://arxiv.org/abs/2412.11985