From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mairittha, Tittaya, Sawanglok, Tanakon, Raden, Panuwit, Buntub, Jirapast, Warunee, Thanapat, Asawachaisuvikrom, Napat, Saiwongin, Thanaphum
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917150402281472
author Mairittha, Tittaya
Sawanglok, Tanakon
Raden, Panuwit
Buntub, Jirapast
Warunee, Thanapat
Asawachaisuvikrom, Napat
Saiwongin, Thanaphum
author_facet Mairittha, Tittaya
Sawanglok, Tanakon
Raden, Panuwit
Buntub, Jirapast
Warunee, Thanapat
Asawachaisuvikrom, Napat
Saiwongin, Thanaphum
contents While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11724
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines
Mairittha, Tittaya
Sawanglok, Tanakon
Raden, Panuwit
Buntub, Jirapast
Warunee, Thanapat
Asawachaisuvikrom, Napat
Saiwongin, Thanaphum
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Software Engineering
While voice-based AI systems have achieved remarkable generative capabilities, their interactions often feel conversationally broken. This paper examines the interactional friction that emerges in modular Speech-to-Speech Retrieval-Augmented Generation (S2S-RAG) pipelines. By analyzing a representative production system, we move beyond simple latency metrics to identify three recurring patterns of conversational breakdown: (1) Temporal Misalignment, where system delays violate user expectations of conversational rhythm; (2) Expressive Flattening, where the loss of paralinguistic cues leads to literal, inappropriate responses; and (3) Repair Rigidity, where architectural gating prevents users from correcting errors in real-time. Through system-level analysis, we demonstrate that these friction points should not be understood as defects or failures, but as structural consequences of a modular design that prioritizes control over fluidity. We conclude that building natural spoken AI is an infrastructure design challenge, requiring a shift from optimizing isolated components to carefully choreographing the seams between them.
title From Signal to Turn: Interactional Friction in Modular Speech-to-Speech Pipelines
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Software Engineering
url https://arxiv.org/abs/2512.11724