SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lou, Haowei, Huang, Chengkai, Paik, Hye-young, Hu, Yongquan, Quigley, Aaron, Hu, Wen, Yao, Lina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915571527843840
author Lou, Haowei
Huang, Chengkai
Paik, Hye-young
Hu, Yongquan
Quigley, Aaron
Hu, Wen
Yao, Lina
author_facet Lou, Haowei
Huang, Chengkai
Paik, Hye-young
Hu, Yongquan
Quigley, Aaron
Hu, Wen
Yao, Lina
contents Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic speech recognition (ASR) and text-to-speech (TTS) technologies, accessible web and mobile infrastructures for users with impaired speech remain limited, hindering the practical adoption of these advances in daily communication. To bridge this gap, we present SpeechAgent, a mobile SpeechAgent designed to facilitate people with speech impairments in everyday communication. The system integrates large language model (LLM)- driven reasoning with advanced speech processing modules, providing adaptive support tailored to diverse impairment types. To ensure real-world practicality, we develop a structured deployment pipeline that enables real-time speech processing on mobile and edge devices, achieving imperceptible latency while maintaining high accuracy and speech quality. Evaluation on real-world impaired speech datasets and edge-device latency profiling confirms that SpeechAgent delivers both effective and user-friendly performance, demonstrating its feasibility for personalized, day-to-day assistive communication.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
Lou, Haowei
Huang, Chengkai
Paik, Hye-young
Hu, Yongquan
Quigley, Aaron
Hu, Wen
Yao, Lina
Systems and Control
Sound
Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic speech recognition (ASR) and text-to-speech (TTS) technologies, accessible web and mobile infrastructures for users with impaired speech remain limited, hindering the practical adoption of these advances in daily communication. To bridge this gap, we present SpeechAgent, a mobile SpeechAgent designed to facilitate people with speech impairments in everyday communication. The system integrates large language model (LLM)- driven reasoning with advanced speech processing modules, providing adaptive support tailored to diverse impairment types. To ensure real-world practicality, we develop a structured deployment pipeline that enables real-time speech processing on mobile and edge devices, achieving imperceptible latency while maintaining high accuracy and speech quality. Evaluation on real-world impaired speech datasets and edge-device latency profiling confirms that SpeechAgent delivers both effective and user-friendly performance, demonstrating its feasibility for personalized, day-to-day assistive communication.
title SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
topic Systems and Control
Sound
url https://arxiv.org/abs/2510.20113