Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912624841588736 |
|---|---|
| author | Elmers, Mikey Inoue, Koji Lala, Divesh Kawahara, Tatsuya |
| author_facet | Elmers, Mikey Inoue, Koji Lala, Divesh Kawahara, Tatsuya |
| contents | Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_07518 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems Elmers, Mikey Inoue, Koji Lala, Divesh Kawahara, Tatsuya Computation and Language Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems. |
| title | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2507.07518 |