Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yiwen, Chakraborti, Monideep, Zhang, Tianyi, Eng, Katelyn, Mohan, Aanchan, Prpa, Mirjana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910888723742720
author Xu, Yiwen
Chakraborti, Monideep
Zhang, Tianyi
Eng, Katelyn
Mohan, Aanchan
Prpa, Mirjana
author_facet Xu, Yiwen
Chakraborti, Monideep
Zhang, Tianyi
Eng, Katelyn
Mohan, Aanchan
Prpa, Mirjana
contents In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication
Xu, Yiwen
Chakraborti, Monideep
Zhang, Tianyi
Eng, Katelyn
Mohan, Aanchan
Prpa, Mirjana
Human-Computer Interaction
Artificial Intelligence
In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity.
title Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2503.17479