Prompt-Guided Turn-Taking Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Inoue, Koji, Elmers, Mikey, Fu, Yahui, Pang, Zi Haur, Lala, Divesh, Ochi, Keiko, Kawahara, Tatsuya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912461572014080
author Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
author_facet Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
contents Turn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study, we propose a novel model that enables turn-taking prediction to be dynamically controlled via textual prompts. This approach allows intuitive and explicit control through instructions such as "faster" or "calmer" adapting dynamically to conversational partners and contexts. The proposed model builds upon a transformer-based voice activity projection (VAP) model, incorporating textual prompt embeddings into both channel-wise transformers and a cross-channel transformer. We evaluated the feasibility of our approach using over 950 hours of human-human spoken dialogue data. Since textual prompt data for the proposed approach was not available in existing datasets, we utilized a large language model (LLM) to generate synthetic prompt sentences. Experimental results demonstrated that the proposed model improved prediction accuracy and effectively varied turn-taking timing behaviors according to the textual prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21191
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt-Guided Turn-Taking Prediction
Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
Computation and Language
Sound
Audio and Speech Processing
Turn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study, we propose a novel model that enables turn-taking prediction to be dynamically controlled via textual prompts. This approach allows intuitive and explicit control through instructions such as "faster" or "calmer" adapting dynamically to conversational partners and contexts. The proposed model builds upon a transformer-based voice activity projection (VAP) model, incorporating textual prompt embeddings into both channel-wise transformers and a cross-channel transformer. We evaluated the feasibility of our approach using over 950 hours of human-human spoken dialogue data. Since textual prompt data for the proposed approach was not available in existing datasets, we utilized a large language model (LLM) to generate synthetic prompt sentences. Experimental results demonstrated that the proposed model improved prediction accuracy and effectively varied turn-taking timing behaviors according to the textual prompts.
title Prompt-Guided Turn-Taking Prediction
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.21191