FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911410111381504 |
|---|---|
| author | Yao, Yiqun Li, Xiang Jiang, Xin Fang, Xuezhi Yu, Naitong Ma, Wenjia Sun, Aixin Wang, Yequan |
| author_facet | Yao, Yiqun Li, Xiang Jiang, Xin Fang, Xuezhi Yu, Naitong Ma, Wenjia Sun, Aixin Wang, Yequan |
| contents | Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution merges multiple channels in each time step, achieving the lowest latency. However, prevailing designs break down the textual monologue sentences for word-level alignment with audio streams, which degrades language modeling abilities. To help address this issue, we introduce "contiguous monologues", which are composed by continuous sentences and "waiting" intervals, mimicking human-like cognitive behavior in dialogs. We find a proper training paradigm to be critical for semantically aligning contiguous monologues with audio. To this end, we develop a "dual" training paradigm that alternates the position of the monologues, either leading or trailing the audio, across different training stages. A combination of our contiguous monologue and dual training strategy is applied in developing FLM-Audio, our 7B spoken dialog chatbot with native full-duplexity. As confirmed by experimental results, FLM-Audio achieves superior response qualities and chatting experiences while requiring significantly less training data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_02521 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training Yao, Yiqun Li, Xiang Jiang, Xin Fang, Xuezhi Yu, Naitong Ma, Wenjia Sun, Aixin Wang, Yequan Sound Artificial Intelligence Computation and Language Full-duplex dialog models aim to listen and speak simultaneously, delivering rapid responses to dynamic user input. Among different solutions to full-duplexity, a native solution merges multiple channels in each time step, achieving the lowest latency. However, prevailing designs break down the textual monologue sentences for word-level alignment with audio streams, which degrades language modeling abilities. To help address this issue, we introduce "contiguous monologues", which are composed by continuous sentences and "waiting" intervals, mimicking human-like cognitive behavior in dialogs. We find a proper training paradigm to be critical for semantically aligning contiguous monologues with audio. To this end, we develop a "dual" training paradigm that alternates the position of the monologues, either leading or trailing the audio, across different training stages. A combination of our contiguous monologue and dual training strategy is applied in developing FLM-Audio, our 7B spoken dialog chatbot with native full-duplexity. As confirmed by experimental results, FLM-Audio achieves superior response qualities and chatting experiences while requiring significantly less training data. |
| title | FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training |
| topic | Sound Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.02521 |