Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Elmers, Mikey, Inoue, Koji, Lala, Divesh, Kawahara, Tatsuya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Analysis and Detection of Differences in Spoken User Behaviors between Autonomous and Wizard-of-Oz Systems
by: Elmers, Mikey, et al.
Published: (2024)
by: Elmers, Mikey, et al.
Published: (2024)
Why Do We Laugh? Annotation and Taxonomy Generation for Laughable Contexts in Spontaneous Text Conversation
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Does the Appearance of Autonomous Conversational Robots Affect User Spoken Behaviors in Real-World Conference Interactions?
by: Pang, Zi Haur, et al.
Published: (2025)
by: Pang, Zi Haur, et al.
Published: (2025)
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Prompt-Guided Turn-Taking Prediction
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference
by: Pang, Zi Haur, et al.
Published: (2024)
by: Pang, Zi Haur, et al.
Published: (2024)
Multilingual Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
by: Inoue, Koji, et al.
Published: (2024)
by: Inoue, Koji, et al.
Published: (2024)
Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Acknowledgment of Emotional States: Generating Validating Responses for Empathetic Dialogue
by: Pang, Zi Haur, et al.
Published: (2024)
by: Pang, Zi Haur, et al.
Published: (2024)
Evaluation of a semi-autonomous attentive listening system with takeover prompting
by: Kawai, Haruki, et al.
Published: (2024)
by: Kawai, Haruki, et al.
Published: (2024)
A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
by: Kato, Kazushi, et al.
Published: (2025)
by: Kato, Kazushi, et al.
Published: (2025)
Enhancing Long-term RAG Chatbots with Psychological Models of Memory Importance and Forgetting
by: Sumida, Ryuichi, et al.
Published: (2024)
by: Sumida, Ryuichi, et al.
Published: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Minority-Aware Satisfaction Estimation in Dialogue Systems via Preference-Adaptive Reinforcement Learning
by: Fu, Yahui, et al.
Published: (2025)
by: Fu, Yahui, et al.
Published: (2025)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
by: Li, Guojian, et al.
Published: (2025)
by: Li, Guojian, et al.
Published: (2025)
Who Speaks Next? Multi-party AI Discussion Leveraging the Systematics of Turn-taking in Murder Mystery Games
by: Nonomura, Ryota, et al.
Published: (2024)
by: Nonomura, Ryota, et al.
Published: (2024)
Enhancing Personality Recognition in Dialogue by Data Augmentation and Heterogeneous Conversational Graph Networks
by: Fu, Yahui, et al.
Published: (2024)
by: Fu, Yahui, et al.
Published: (2024)
JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems
by: Yang, Guangzhao, et al.
Published: (2026)
by: Yang, Guangzhao, et al.
Published: (2026)
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
by: Zhu, Han, et al.
Published: (2025)
by: Zhu, Han, et al.
Published: (2025)
Are LLMs Robust for Spoken Dialogues?
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
StyEmp: Stylizing Empathetic Response Generation via Multi-Grained Prefix Encoder and Personality Reinforcement
by: Fu, Yahui, et al.
Published: (2024)
by: Fu, Yahui, et al.
Published: (2024)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
by: Lee, Jonggeun, et al.
Published: (2026)
by: Lee, Jonggeun, et al.
Published: (2026)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Human Latency Conversational Turns for Spoken Avatar Systems
by: Jacoby, Derek, et al.
Published: (2024)
by: Jacoby, Derek, et al.
Published: (2024)
Contrastive Speaker-Aware Learning for Multi-party Dialogue Generation with LLMs
by: Sun, Tianyu, et al.
Published: (2025)
by: Sun, Tianyu, et al.
Published: (2025)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
by: Matsuura, Ryuki, et al.
Published: (2025)
by: Matsuura, Ryuki, et al.
Published: (2025)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
by: Tu, Wenming, et al.
Published: (2025)
by: Tu, Wenming, et al.
Published: (2025)
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
by: Deng, Yayue, et al.
Published: (2025)
by: Deng, Yayue, et al.
Published: (2025)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
by: Yoon, Jaeseok, et al.
Published: (2023)
by: Yoon, Jaeseok, et al.
Published: (2023)
Voice Activity Projection Model with Multimodal Encoders
by: Saga, Takeshi, et al.
Published: (2025)
by: Saga, Takeshi, et al.
Published: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
by: Nakata, Wataru, et al.
Published: (2024)
by: Nakata, Wataru, et al.
Published: (2024)
Towards a Japanese Full-duplex Spoken Dialogue System
by: Ohashi, Atsumoto, et al.
Published: (2025)
by: Ohashi, Atsumoto, et al.
Published: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Similar Items
-
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
by: Inoue, Koji, et al.
Published: (2025) -
Analysis and Detection of Differences in Spoken User Behaviors between Autonomous and Wizard-of-Oz Systems
by: Elmers, Mikey, et al.
Published: (2024) -
Why Do We Laugh? Annotation and Taxonomy Generation for Laughable Contexts in Spontaneous Text Conversation
by: Inoue, Koji, et al.
Published: (2025) -
Does the Appearance of Autonomous Conversational Robots Affect User Spoken Behaviors in Real-World Conference Interactions?
by: Pang, Zi Haur, et al.
Published: (2025) -
An Analysis of User Behaviors for Objectively Evaluating Spoken Dialogue Systems
by: Inoue, Koji, et al.
Published: (2024)