Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908498276646912 |
|---|---|
| author | Allbert, Rumi Yazdani, Nima Ansari, Ali Mahajan, Aruj Afsharrad, Amirhossein Mousavi, Seyed Shahabeddin |
| author_facet | Allbert, Rumi Yazdani, Nima Ansari, Ali Mahajan, Aruj Afsharrad, Amirhossein Mousavi, Seyed Shahabeddin |
| contents | Voice-based conversational AI systems increasingly rely on cascaded architectures that combine speech-to-text (STT), large language models (LLMs), and text-to-speech (TTS) components. We present a large-scale empirical comparison of STT x LLM x TTS stacks using data sampled from over 300,000 AI-conducted job interviews. We used an LLM-as-a-Judge automated evaluation framework to assess conversational quality, technical accuracy, and skill assessment capabilities. Our analysis of five production configurations reveals that a stack combining Google's STT, GPT-4.1, and Cartesia's TTS outperforms alternatives in both objective quality metrics and user satisfaction scores. Surprisingly, we find that objective quality metrics correlate weakly with user satisfaction scores, suggesting that user experience in voice-based AI systems depends on factors beyond technical performance. Our findings provide practical guidance for selecting components in multimodal conversations and contribute a validated evaluation methodology for human-AI interactions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_16835 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Allbert, Rumi Yazdani, Nima Ansari, Ali Mahajan, Aruj Afsharrad, Amirhossein Mousavi, Seyed Shahabeddin Audio and Speech Processing Computation and Language Voice-based conversational AI systems increasingly rely on cascaded architectures that combine speech-to-text (STT), large language models (LLMs), and text-to-speech (TTS) components. We present a large-scale empirical comparison of STT x LLM x TTS stacks using data sampled from over 300,000 AI-conducted job interviews. We used an LLM-as-a-Judge automated evaluation framework to assess conversational quality, technical accuracy, and skill assessment capabilities. Our analysis of five production configurations reveals that a stack combining Google's STT, GPT-4.1, and Cartesia's TTS outperforms alternatives in both objective quality metrics and user satisfaction scores. Surprisingly, we find that objective quality metrics correlate weakly with user satisfaction scores, suggesting that user experience in voice-based AI systems depends on factors beyond technical performance. Our findings provide practical guidance for selecting components in multimodal conversations and contribute a validated evaluation methodology for human-AI interactions. |
| title | Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems |
| topic | Audio and Speech Processing Computation and Language |
| url | https://arxiv.org/abs/2507.16835 |