ChipChat: Low-Latency Cascaded Conversational Agent in MLX
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918132863467520 |
|---|---|
| author | Likhomanenko, Tatiana Carlson, Luke Bai, Richard He Gu, Zijin Tran, Han Aldeneh, Zakaria Zhang, Yizhe Zhang, Ruixiang Zheng, Huangjie Jaitly, Navdeep |
| author_facet | Likhomanenko, Tatiana Carlson, Luke Bai, Richard He Gu, Zijin Tran, Han Aldeneh, Zakaria Zhang, Yizhe Zhang, Ruixiang Zheng, Huangjie Jaitly, Navdeep |
| contents | The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_00078 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ChipChat: Low-Latency Cascaded Conversational Agent in MLX Likhomanenko, Tatiana Carlson, Luke Bai, Richard He Gu, Zijin Tran, Han Aldeneh, Zakaria Zhang, Yizhe Zhang, Ruixiang Zheng, Huangjie Jaitly, Navdeep Audio and Speech Processing Computation and Language Machine Learning Sound The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents. |
| title | ChipChat: Low-Latency Cascaded Conversational Agent in MLX |
| topic | Audio and Speech Processing Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2509.00078 |