ChipChat: Low-Latency Cascaded Conversational Agent in MLX

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Likhomanenko, Tatiana, Carlson, Luke, Bai, Richard He, Gu, Zijin, Tran, Han, Aldeneh, Zakaria, Zhang, Yizhe, Zhang, Ruixiang, Zheng, Huangjie, Jaitly, Navdeep
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918132863467520
author Likhomanenko, Tatiana
Carlson, Luke
Bai, Richard He
Gu, Zijin
Tran, Han
Aldeneh, Zakaria
Zhang, Yizhe
Zhang, Ruixiang
Zheng, Huangjie
Jaitly, Navdeep
author_facet Likhomanenko, Tatiana
Carlson, Luke
Bai, Richard He
Gu, Zijin
Tran, Han
Aldeneh, Zakaria
Zhang, Yizhe
Zhang, Ruixiang
Zheng, Huangjie
Jaitly, Navdeep
contents The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChipChat: Low-Latency Cascaded Conversational Agent in MLX
Likhomanenko, Tatiana
Carlson, Luke
Bai, Richard He
Gu, Zijin
Tran, Han
Aldeneh, Zakaria
Zhang, Yizhe
Zhang, Ruixiang
Zheng, Huangjie
Jaitly, Navdeep
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.
title ChipChat: Low-Latency Cascaded Conversational Agent in MLX
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2509.00078