SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gautam, Sushant, Sarkhoosh, Mehdi Houshmand, Held, Jan, Midoglu, Cise, Cioppa, Anthony, Giancola, Silvio, Thambawita, Vajira, Riegler, Michael A., Halvorsen, Pål, Shah, Mubarak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914001220272128
author Gautam, Sushant
Sarkhoosh, Mehdi Houshmand
Held, Jan
Midoglu, Cise
Cioppa, Anthony
Giancola, Silvio
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
Shah, Mubarak
author_facet Gautam, Sushant
Sarkhoosh, Mehdi Houshmand
Held, Jan
Midoglu, Cise
Cioppa, Anthony
Giancola, Silvio
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
Shah, Mubarak
contents The application of Automatic Speech Recognition (ASR) technology in soccer offers numerous opportunities for sports analytics. Specifically, extracting audio commentaries with ASR provides valuable insights into the events of the game, and opens the door to several downstream applications such as automatic highlight generation. This paper presents SoccerNet-Echoes, an augmentation of the SoccerNet dataset with automatically generated transcriptions of audio commentaries from soccer game broadcasts, enhancing video content with rich layers of textual information derived from the game audio using ASR. These textual commentaries, generated using the Whisper model and translated with Google Translate, extend the usefulness of the SoccerNet dataset in diverse applications such as enhanced action spotting, automatic caption generation, and game summarization. By incorporating textual data alongside visual and auditory content, SoccerNet-Echoes aims to serve as a comprehensive resource for the development of algorithms specialized in capturing the dynamics of soccer games. We detail the methods involved in the curation of this dataset and the integration of ASR. We also highlight the implications of a multimodal approach in sports analytics, and how the enriched dataset can support diverse applications, thus broadening the scope of research and development in the field of sports analytics.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
Gautam, Sushant
Sarkhoosh, Mehdi Houshmand
Held, Jan
Midoglu, Cise
Cioppa, Anthony
Giancola, Silvio
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
Shah, Mubarak
Sound
Information Retrieval
Machine Learning
Multimedia
Audio and Speech Processing
I.2.7; I.7
The application of Automatic Speech Recognition (ASR) technology in soccer offers numerous opportunities for sports analytics. Specifically, extracting audio commentaries with ASR provides valuable insights into the events of the game, and opens the door to several downstream applications such as automatic highlight generation. This paper presents SoccerNet-Echoes, an augmentation of the SoccerNet dataset with automatically generated transcriptions of audio commentaries from soccer game broadcasts, enhancing video content with rich layers of textual information derived from the game audio using ASR. These textual commentaries, generated using the Whisper model and translated with Google Translate, extend the usefulness of the SoccerNet dataset in diverse applications such as enhanced action spotting, automatic caption generation, and game summarization. By incorporating textual data alongside visual and auditory content, SoccerNet-Echoes aims to serve as a comprehensive resource for the development of algorithms specialized in capturing the dynamics of soccer games. We detail the methods involved in the curation of this dataset and the integration of ASR. We also highlight the implications of a multimodal approach in sports analytics, and how the enriched dataset can support diverse applications, thus broadening the scope of research and development in the field of sports analytics.
title SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
topic Sound
Information Retrieval
Machine Learning
Multimedia
Audio and Speech Processing
I.2.7; I.7
url https://arxiv.org/abs/2405.07354