Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Inoue, Koji, Elmers, Mikey, Fu, Yahui, Pang, Zi Haur, Mori, Taiga, Lala, Divesh, Ochi, Keiko, Kawahara, Tatsuya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917148405792768
author Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Mori, Taiga
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
author_facet Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Mori, Taiga
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
contents We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame level, jointly trained with auxiliary tasks on approximately 300 hours of dyadic conversations. Across all three languages, the multilingual model matches or surpasses monolingual baselines, indicating that it learns both language-universal cues and language-specific timing patterns. Zero-shot transfer with two-language training remains limited, underscoring substantive cross-lingual differences. Perturbation analyses reveal distinct cue usage: Japanese relies more on short-term linguistic information, whereas English and Chinese are more sensitive to silence duration and prosodic variation; multilingual training encourages shared yet adaptable representations and reduces overreliance on pitch in Chinese. A context-length study further shows that Japanese is relatively robust to shorter contexts, while Chinese benefits markedly from longer contexts. Finally, we integrate the trained model into a real-time processing software, demonstrating CPU-only inference. Together, these findings provide a unified model and empirical evidence for how backchannel timing differs across languages, informing the design of more natural, culturally-aware spoken dialogue systems.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
Inoue, Koji
Elmers, Mikey
Fu, Yahui
Pang, Zi Haur
Mori, Taiga
Lala, Divesh
Ochi, Keiko
Kawahara, Tatsuya
Computation and Language
Human-Computer Interaction
Sound
We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame level, jointly trained with auxiliary tasks on approximately 300 hours of dyadic conversations. Across all three languages, the multilingual model matches or surpasses monolingual baselines, indicating that it learns both language-universal cues and language-specific timing patterns. Zero-shot transfer with two-language training remains limited, underscoring substantive cross-lingual differences. Perturbation analyses reveal distinct cue usage: Japanese relies more on short-term linguistic information, whereas English and Chinese are more sensitive to silence duration and prosodic variation; multilingual training encourages shared yet adaptable representations and reduces overreliance on pitch in Chinese. A context-length study further shows that Japanese is relatively robust to shorter contexts, while Chinese benefits markedly from longer contexts. Finally, we integrate the trained model into a real-time processing software, demonstrating CPU-only inference. Together, these findings provide a unified model and empirical evidence for how backchannel timing differs across languages, informing the design of more natural, culturally-aware spoken dialogue systems.
title Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
topic Computation and Language
Human-Computer Interaction
Sound
url https://arxiv.org/abs/2512.14085