Emotion-Aware Speech Generation with Character-Specific Voices for Comics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Zhiwen, Liang, Jinhua, Zhang, Huan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908546398945280
author Qian, Zhiwen
Liang, Jinhua
Zhang, Huan
author_facet Qian, Zhiwen
Liang, Jinhua
Zhang, Huan
contents This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emotion-Aware Speech Generation with Character-Specific Voices for Comics
Qian, Zhiwen
Liang, Jinhua
Zhang, Huan
Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience.
title Emotion-Aware Speech Generation with Character-Specific Voices for Comics
topic Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2509.15253