From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917150402281472 |
|---|---|
| author | Mairittha, Tittaya Sawanglok, Tanakon Raden, Panuwit Buntub, Jirapast Warunee, Thanapat Asawachaisuvikrom, Napat Saiwongin, Thanaphum |
| author_facet | Mairittha, Tittaya Sawanglok, Tanakon Raden, Panuwit Buntub, Jirapast Warunee, Thanapat Asawachaisuvikrom, Napat Saiwongin, Thanaphum |
| contents | While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_11724 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines Mairittha, Tittaya Sawanglok, Tanakon Raden, Panuwit Buntub, Jirapast Warunee, Thanapat Asawachaisuvikrom, Napat Saiwongin, Thanaphum Human-Computer Interaction Artificial Intelligence Computation and Language Software Engineering While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them. |
| title | From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines |
| topic | Human-Computer Interaction Artificial Intelligence Computation and Language Software Engineering |
| url | https://arxiv.org/abs/2512.11724 |