Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910888723742720 |
|---|---|
| author | Xu, Yiwen Chakraborti, Monideep Zhang, Tianyi Eng, Katelyn Mohan, Aanchan Prpa, Mirjana |
| author_facet | Xu, Yiwen Chakraborti, Monideep Zhang, Tianyi Eng, Katelyn Mohan, Aanchan Prpa, Mirjana |
| contents | In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_17479 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication Xu, Yiwen Chakraborti, Monideep Zhang, Tianyi Eng, Katelyn Mohan, Aanchan Prpa, Mirjana Human-Computer Interaction Artificial Intelligence In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity. |
| title | Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication |
| topic | Human-Computer Interaction Artificial Intelligence |
| url | https://arxiv.org/abs/2503.17479 |