EchoVoices: Preserving Generational Voices and Memories for Seniors and Children

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Haiying, Liu, Haoze, Li, Mingshi, Cai, Siyu, Zheng, Guangxuan, Jia, Yuhuang, Zhao, Jinghua, Qin, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912493356449792
author Xu, Haiying
Liu, Haoze
Li, Mingshi
Cai, Siyu
Zheng, Guangxuan
Jia, Yuhuang
Zhao, Jinghua
Qin, Yong
author_facet Xu, Haiying
Liu, Haoze
Li, Mingshi
Cai, Siyu
Zheng, Guangxuan
Jia, Yuhuang
Zhao, Jinghua
Qin, Yong
contents Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics possess distinct vocal characteristics, linguistic styles, and interaction patterns that challenge conventional ASR, TTS, and LLM systems. To address this, we introduce EchoVoices, an end-to-end digital human pipeline dedicated to creating persistent digital personas for seniors and children, ensuring their voices and memories are preserved for future generations. Our system integrates three core innovations: a k-NN-enhanced Whisper model for robust speech recognition of atypical speech; an age-adaptive VITS model for high-fidelity, speaker-aware speech synthesis; and an LLM-driven agent that automatically generates persona cards and leverages a RAG-based memory system for conversational consistency. Our experiments, conducted on the SeniorTalk and ChildMandarin datasets, demonstrate significant improvements in recognition accuracy, synthesis quality, and speaker similarity. EchoVoices provides a comprehensive framework for preserving generational voices, offering a new means of intergenerational connection and the creation of lasting digital legacies.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15221
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
Xu, Haiying
Liu, Haoze
Li, Mingshi
Cai, Siyu
Zheng, Guangxuan
Jia, Yuhuang
Zhao, Jinghua
Qin, Yong
Sound
Audio and Speech Processing
Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics possess distinct vocal characteristics, linguistic styles, and interaction patterns that challenge conventional ASR, TTS, and LLM systems. To address this, we introduce EchoVoices, an end-to-end digital human pipeline dedicated to creating persistent digital personas for seniors and children, ensuring their voices and memories are preserved for future generations. Our system integrates three core innovations: a k-NN-enhanced Whisper model for robust speech recognition of atypical speech; an age-adaptive VITS model for high-fidelity, speaker-aware speech synthesis; and an LLM-driven agent that automatically generates persona cards and leverages a RAG-based memory system for conversational consistency. Our experiments, conducted on the SeniorTalk and ChildMandarin datasets, demonstrate significant improvements in recognition accuracy, synthesis quality, and speaker similarity. EchoVoices provides a comprehensive framework for preserving generational voices, offering a new means of intergenerational connection and the creation of lasting digital legacies.
title EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.15221