Saved in:
Bibliographic Details
Main Authors: Xiao, Cihan, Liang, Ruixing, Zhang, Xiangyu, Tiryaki, Mehmet Emre, Bae, Veronica, Shankar, Lavanya, Yang, Rong, Poon, Ethan, Dupoux, Emmanuel, Khudanpur, Sanjeev, Perera, Leibny Paola Garcia
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.00267
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909645664157696
author Xiao, Cihan
Liang, Ruixing
Zhang, Xiangyu
Tiryaki, Mehmet Emre
Bae, Veronica
Shankar, Lavanya
Yang, Rong
Poon, Ethan
Dupoux, Emmanuel
Khudanpur, Sanjeev
Perera, Leibny Paola Garcia
author_facet Xiao, Cihan
Liang, Ruixing
Zhang, Xiangyu
Tiryaki, Mehmet Emre
Bae, Veronica
Shankar, Lavanya
Yang, Rong
Poon, Ethan
Dupoux, Emmanuel
Khudanpur, Sanjeev
Perera, Leibny Paola Garcia
contents The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00267
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CASPER: A Large Scale Spontaneous Speech Dataset
Xiao, Cihan
Liang, Ruixing
Zhang, Xiangyu
Tiryaki, Mehmet Emre
Bae, Veronica
Shankar, Lavanya
Yang, Rong
Poon, Ethan
Dupoux, Emmanuel
Khudanpur, Sanjeev
Perera, Leibny Paola Garcia
Computation and Language
Sound
Audio and Speech Processing
The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.
title CASPER: A Large Scale Spontaneous Speech Dataset
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00267