Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Ziqi, Chen, Yiting, Xu, Jiacheng, Xie, Liufei, Wang, Yuchen, Yang, Zhenchuan, Bai, Bingsong, Gao, Yangsheng, Zhou, Wenjiang, Zhao, Weifeng, Zhou, Ruohua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914047202426880
author Dai, Ziqi
Chen, Yiting
Xu, Jiacheng
Xie, Liufei
Wang, Yuchen
Yang, Zhenchuan
Bai, Bingsong
Gao, Yangsheng
Zhou, Wenjiang
Zhao, Weifeng
Zhou, Ruohua
author_facet Dai, Ziqi
Chen, Yiting
Xu, Jiacheng
Xie, Liufei
Wang, Yuchen
Yang, Zhenchuan
Bai, Bingsong
Gao, Yangsheng
Zhou, Wenjiang
Zhao, Weifeng
Zhou, Ruohua
contents The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP models, whereas character voice timbre selection still relies on manual effort. Speech synthesis uses either manual dubbing or text-to-speech (TTS). While TTS boosts efficiency, it struggles with emotional expression, intonation control, and contextual scene adaptation. To address these challenges, we propose DeepDubbing, an end-to-end automated system for multi-participant audiobook production. The system comprises two main components: a Text-to-Timbre (TTT) model and a Context-Aware Instruct-TTS (CA-Instruct-TTS) model. The TTT model generates role-specific timbre embeddings conditioned on text descriptions. The CA-Instruct-TTS model synthesizes expressive speech by analyzing contextual dialogue and incorporating fine-grained emotional instructions. This system enables the automated generation of multi-participant audiobooks with both timbre-matched character voices and emotionally expressive narration, offering a novel solution for audiobook production.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
Dai, Ziqi
Chen, Yiting
Xu, Jiacheng
Xie, Liufei
Wang, Yuchen
Yang, Zhenchuan
Bai, Bingsong
Gao, Yangsheng
Zhou, Wenjiang
Zhao, Weifeng
Zhou, Ruohua
Audio and Speech Processing
I.2.7
The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP models, whereas character voice timbre selection still relies on manual effort. Speech synthesis uses either manual dubbing or text-to-speech (TTS). While TTS boosts efficiency, it struggles with emotional expression, intonation control, and contextual scene adaptation. To address these challenges, we propose DeepDubbing, an end-to-end automated system for multi-participant audiobook production. The system comprises two main components: a Text-to-Timbre (TTT) model and a Context-Aware Instruct-TTS (CA-Instruct-TTS) model. The TTT model generates role-specific timbre embeddings conditioned on text descriptions. The CA-Instruct-TTS model synthesizes expressive speech by analyzing contextual dialogue and incorporating fine-grained emotional instructions. This system enables the automated generation of multi-participant audiobooks with both timbre-matched character voices and emotionally expressive narration, offering a novel solution for audiobook production.
title Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
topic Audio and Speech Processing
I.2.7
url https://arxiv.org/abs/2509.15845