A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Runchuan, Zhou, Yixuan, Yu, Renjie, Lin, Zijian, Li, Kehan, Li, Xiang, Liu, Xin, Zeng, Guoyang, Wu, Zhiyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909988149002240
author Ye, Runchuan
Zhou, Yixuan
Yu, Renjie
Lin, Zijian
Li, Kehan
Li, Xiang
Liu, Xin
Zeng, Guoyang
Wu, Zhiyong
author_facet Ye, Runchuan
Zhou, Yixuan
Yu, Renjie
Lin, Zijian
Li, Kehan
Li, Xiang
Liu, Xin
Zeng, Guoyang
Wu, Zhiyong
contents Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional intelligence. Existing methods for NVs acquisition are either costly and unscalable (relying on manual annotation/recording) or unnatural (relying on rule-based synthesis). To address these limitations, we propose a highly scalable automatic annotation framework to label non-verbal phenomena from natural speech, which is low-cost, easily extendable, and inherently diverse and natural. This framework leverages a unified detection model to accurately identify NVs in natural speech and integrates them with transcripts via temporal-semantic alignment method. Using this framework, we created and released \textbf{NonVerbalSpeech-38K}, a diverse, real-world dataset featuring 38,718 samples across 10 NV categories collected from in-the-wild media. Experimental results demonstrate that our dataset provides superior controllability for NVs generation and achieves comparable performance for NVs understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05385
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
Ye, Runchuan
Zhou, Yixuan
Yu, Renjie
Lin, Zijian
Li, Kehan
Li, Xiang
Liu, Xin
Zeng, Guoyang
Wu, Zhiyong
Sound
Audio and Speech Processing
Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional intelligence. Existing methods for NVs acquisition are either costly and unscalable (relying on manual annotation/recording) or unnatural (relying on rule-based synthesis). To address these limitations, we propose a highly scalable automatic annotation framework to label non-verbal phenomena from natural speech, which is low-cost, easily extendable, and inherently diverse and natural. This framework leverages a unified detection model to accurately identify NVs in natural speech and integrates them with transcripts via temporal-semantic alignment method. Using this framework, we created and released \textbf{NonVerbalSpeech-38K}, a diverse, real-world dataset featuring 38,718 samples across 10 NV categories collected from in-the-wild media. Experimental results demonstrate that our dataset provides superior controllability for NVs generation and achieves comparable performance for NVs understanding.
title A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.05385