Saved in:
Bibliographic Details
Main Authors: Brade, Stephen, Anderson, Sam, Kumar, Rithesh, Jin, Zeyu, Truong, Anh
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.05106
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909569295319040
author Brade, Stephen
Anderson, Sam
Kumar, Rithesh
Jin, Zeyu
Truong, Anh
author_facet Brade, Stephen
Anderson, Sam
Kumar, Rithesh
Jin, Zeyu
Truong, Anh
contents Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents, many struggle with unintuitive or overly granular TTS interfaces. We propose simplifying TTS generation by allowing users to specify high-level context alongside their script. Our Wizard-of-Oz system, SpeakEasy, leverages user-provided context to inform and influence TTS output, enabling iterative refinement with high-level feedback. This approach was informed by two 8-subject formative studies: one examining content creators' experiences with TTS, and the other drawing on effective strategies from voice actors. Our evaluation shows that participants using SpeakEasy were more successful in generating performances matching their personal standards, without requiring significantly more effort than leading industry interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
Brade, Stephen
Anderson, Sam
Kumar, Rithesh
Jin, Zeyu
Truong, Anh
Human-Computer Interaction
Artificial Intelligence
Machine Learning
Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents, many struggle with unintuitive or overly granular TTS interfaces. We propose simplifying TTS generation by allowing users to specify high-level context alongside their script. Our Wizard-of-Oz system, SpeakEasy, leverages user-provided context to inform and influence TTS output, enabling iterative refinement with high-level feedback. This approach was informed by two 8-subject formative studies: one examining content creators' experiences with TTS, and the other drawing on effective strategies from voice actors. Our evaluation shows that participants using SpeakEasy were more successful in generating performances matching their personal standards, without requiring significantly more effort than leading industry interfaces.
title SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
topic Human-Computer Interaction
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.05106