Summary of the NOTSOFAR-1 Challenge: Highlights and Learnings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abramovski, Igor, Vinnikov, Alon, Shaer, Shalev, Kanda, Naoyuki, Wang, Xiaofei, Ivry, Amir, Krupka, Eyal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912265384493056
author Abramovski, Igor
Vinnikov, Alon
Shaer, Shalev
Kanda, Naoyuki
Wang, Xiaofei
Ivry, Amir
Krupka, Eyal
author_facet Abramovski, Igor
Vinnikov, Alon
Shaer, Shalev
Kanda, Naoyuki
Wang, Xiaofei
Ivry, Amir
Krupka, Eyal
contents The first Natural Office Talkers in Settings of Far-field Audio Recordings (NOTSOFAR-1) Challenge is a pivotal initiative that sets new benchmarks by offering datasets more representative of the needs of real-world business applications than those previously available. The challenge provides a unique combination of 280 recorded meetings across 30 diverse environments, capturing real-world acoustic conditions and conversational dynamics, and a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. In this paper, we provide an overview of the systems submitted to the challenge and analyze the top-performing approaches, hypothesizing the factors behind their success. Additionally, we highlight promising directions left unexplored by participants. By presenting key findings and actionable insights, this work aims to drive further innovation and progress in DASR research and applications.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Summary of the NOTSOFAR-1 Challenge: Highlights and Learnings
Abramovski, Igor
Vinnikov, Alon
Shaer, Shalev
Kanda, Naoyuki
Wang, Xiaofei
Ivry, Amir
Krupka, Eyal
Sound
Machine Learning
Audio and Speech Processing
The first Natural Office Talkers in Settings of Far-field Audio Recordings (NOTSOFAR-1) Challenge is a pivotal initiative that sets new benchmarks by offering datasets more representative of the needs of real-world business applications than those previously available. The challenge provides a unique combination of 280 recorded meetings across 30 diverse environments, capturing real-world acoustic conditions and conversational dynamics, and a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. In this paper, we provide an overview of the systems submitted to the challenge and analyze the top-performing approaches, hypothesizing the factors behind their success. Additionally, we highlight promising directions left unexplored by participants. By presenting key findings and actionable insights, this work aims to drive further innovation and progress in DASR research and applications.
title Summary of the NOTSOFAR-1 Challenge: Highlights and Learnings
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2501.17304