Emotion-Aware Speech Generation with Character-Specific Voices for Comics
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908546398945280 |
|---|---|
| author | Qian, Zhiwen Liang, Jinhua Zhang, Huan |
| author_facet | Qian, Zhiwen Liang, Jinhua Zhang, Huan |
| contents | This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15253 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Emotion-Aware Speech Generation with Character-Specific Voices for Comics Qian, Zhiwen Liang, Jinhua Zhang, Huan Sound Artificial Intelligence Multimedia Audio and Speech Processing This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience. |
| title | Emotion-Aware Speech Generation with Character-Specific Voices for Comics |
| topic | Sound Artificial Intelligence Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.15253 |