EmojiVoice: Towards long-term controllable expressivity in robot speech
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913965586513920 |
|---|---|
| author | Tuttösí, Paige Mehta, Shivam Syvenky, Zachary Burkanova, Bermet Henter, Gustav Eje Lim, Angelica |
| author_facet | Tuttösí, Paige Mehta, Shivam Syvenky, Zachary Burkanova, Bermet Henter, Gustav Eje Lim, Angelica |
| contents | Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with ``expressive'' joyful voices, they lack this long-term variation found in human speech. Foundation model text-to-speech systems are beginning to mimic the expressivity in human speech, but they are difficult to deploy offline on robots. We present EmojiVoice, a free, customizable text-to-speech (TTS) toolkit that allows social roboticists to build temporally variable, expressive speech on social robots. We introduce emoji-prompting to allow fine-grained control of expressivity on a phase level and use the lightweight Matcha-TTS backbone to generate speech in real-time. We explore three case studies: (1) a scripted conversation with a robot assistant, (2) a storytelling robot, and (3) an autonomous speech-to-speech interactive agent. We found that using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task, but expressive voice was not preferred in the assistant use case. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_15085 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EmojiVoice: Towards long-term controllable expressivity in robot speech Tuttösí, Paige Mehta, Shivam Syvenky, Zachary Burkanova, Bermet Henter, Gustav Eje Lim, Angelica Robotics Human-Computer Interaction Humans vary their expressivity when speaking for extended periods to maintain engagement with their listener. Although social robots tend to be deployed with ``expressive'' joyful voices, they lack this long-term variation found in human speech. Foundation model text-to-speech systems are beginning to mimic the expressivity in human speech, but they are difficult to deploy offline on robots. We present EmojiVoice, a free, customizable text-to-speech (TTS) toolkit that allows social roboticists to build temporally variable, expressive speech on social robots. We introduce emoji-prompting to allow fine-grained control of expressivity on a phase level and use the lightweight Matcha-TTS backbone to generate speech in real-time. We explore three case studies: (1) a scripted conversation with a robot assistant, (2) a storytelling robot, and (3) an autonomous speech-to-speech interactive agent. We found that using varied emoji prompting improved the perception and expressivity of speech over a long period in a storytelling task, but expressive voice was not preferred in the assistant use case. |
| title | EmojiVoice: Towards long-term controllable expressivity in robot speech |
| topic | Robotics Human-Computer Interaction |
| url | https://arxiv.org/abs/2506.15085 |