Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fürst, Darren, Steindl, Sebastian, Schäfer, Ulrich
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918474246258688
author Fürst, Darren
Steindl, Sebastian
Schäfer, Ulrich
author_facet Fürst, Darren
Steindl, Sebastian
Schäfer, Ulrich
contents Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable caseloads. We test a hierarchical approach to SSD classification on the granular multi-task SLPHelmUltraSuitePlus benchmark. We propose a cascading approach from binary classification to type, and symptom classification. By fine-tuning Speech Representation Models (SRM), and using targeted data augmentation we mitigate biases found by previous works, and improve upon all clinical tasks in the benchmark. We also treat Automatic Speech Recognition (ASR) with our data augmentation approach. Our results demonstrate that SRM consistently outperform the LLM-based state-of-the-art across all evaluated tasks by a large margin. We publish our models and code to foster future research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26568
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal LLMs are not all you need for Pediatric Speech Language Pathology
Fürst, Darren
Steindl, Sebastian
Schäfer, Ulrich
Computation and Language
Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable caseloads. We test a hierarchical approach to SSD classification on the granular multi-task SLPHelmUltraSuitePlus benchmark. We propose a cascading approach from binary classification to type, and symptom classification. By fine-tuning Speech Representation Models (SRM), and using targeted data augmentation we mitigate biases found by previous works, and improve upon all clinical tasks in the benchmark. We also treat Automatic Speech Recognition (ASR) with our data augmentation approach. Our results demonstrate that SRM consistently outperform the LLM-based state-of-the-art across all evaluated tasks by a large margin. We publish our models and code to foster future research.
title Multimodal LLMs are not all you need for Pediatric Speech Language Pathology
topic Computation and Language
url https://arxiv.org/abs/2604.26568