Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yifan, Liu, Rui, Ren, Yi, Yin, Xiang, Li, Haizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913845734277120
author Hu, Yifan
Liu, Rui
Ren, Yi
Yin, Xiang
Li, Haizhou
author_facet Hu, Yifan
Liu, Rui
Ren, Yi
Yin, Xiang
Li, Haizhou
contents Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
Hu, Yifan
Liu, Rui
Ren, Yi
Yin, Xiang
Li, Haizhou
Sound
Audio and Speech Processing
Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding. To address the above issues, we present Chain-Talker, a three-stage framework mimicking human cognition: Emotion Understanding derives context-aware emotion descriptors from dialogue history; Semantic Understanding generates compact semantic codes via serialized prediction; and Empathetic Rendering synthesizes expressive speech by integrating both components. To support emotion modeling, we develop CSS-EmCap, an LLM-driven automated pipeline for generating precise conversational speech emotion captions. Experiments on three benchmark datasets demonstrate that Chain-Talker produces more expressive and empathetic speech than existing methods, with CSS-EmCap contributing to reliable emotion modeling. The code and demos are available at: https://github.com/AI-S2-Lab/Chain-Talker.
title Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.12597