Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Teleki, Maria, Janjur, Sai, Liu, Haoran, Grabner, Oliver, Verma, Ketan, Docog, Thomas, Dong, Xiangjue, Shi, Lingfeng, Wang, Cong, Birkelbach, Stephanie, Kim, Jason, Zhang, Yin, Székely, Éva, Caverlee, James
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918370478129152
author Teleki, Maria
Janjur, Sai
Liu, Haoran
Grabner, Oliver
Verma, Ketan
Docog, Thomas
Dong, Xiangjue
Shi, Lingfeng
Wang, Cong
Birkelbach, Stephanie
Kim, Jason
Zhang, Yin
Székely, Éva
Caverlee, James
author_facet Teleki, Maria
Janjur, Sai
Liu, Haoran
Grabner, Oliver
Verma, Ketan
Docog, Thomas
Dong, Xiangjue
Shi, Lingfeng
Wang, Cong
Birkelbach, Stephanie
Kim, Jason
Zhang, Yin
Székely, Éva
Caverlee, James
contents LLMs serve as the backbone in SpeechLLMs, yet their behavior on spontaneous conversational input remains poorly understood. Conversational speech contains pervasive disfluencies -- interjections, edits, and parentheticals -- that are rare in the written corpora used for pre-training. Because gold disfluency removal is a deletion-only task, it serves as a controlled probe to determine whether a model performs faithful structural repair or biased reinterpretation. Using the DRES evaluation framework, we evaluate proprietary and open-source LLMs across architectures and scales. We show that model performance clusters into stable precision-recall regimes reflecting distinct editing policies. Notably, reasoning models systematically over-delete fluent content, revealing a bias toward semantic abstraction over structural fidelity. While fine-tuning achieves SOTA results, it harms generalization. Our findings demonstrate that robustness to speech is shaped by specific training objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
Teleki, Maria
Janjur, Sai
Liu, Haoran
Grabner, Oliver
Verma, Ketan
Docog, Thomas
Dong, Xiangjue
Shi, Lingfeng
Wang, Cong
Birkelbach, Stephanie
Kim, Jason
Zhang, Yin
Székely, Éva
Caverlee, James
Computation and Language
Artificial Intelligence
Audio and Speech Processing
LLMs serve as the backbone in SpeechLLMs, yet their behavior on spontaneous conversational input remains poorly understood. Conversational speech contains pervasive disfluencies -- interjections, edits, and parentheticals -- that are rare in the written corpora used for pre-training. Because gold disfluency removal is a deletion-only task, it serves as a controlled probe to determine whether a model performs faithful structural repair or biased reinterpretation. Using the DRES evaluation framework, we evaluate proprietary and open-source LLMs across architectures and scales. We show that model performance clusters into stable precision-recall regimes reflecting distinct editing policies. Notably, reasoning models systematically over-delete fluent content, revealing a bias toward semantic abstraction over structural fidelity. While fine-tuning achieves SOTA results, it harms generalization. Our findings demonstrate that robustness to speech is shaped by specific training objectives.
title Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.20321