Speech to Speech Synthesis for Voice Impersonation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Johnson, Bjorn, Levy, Jared
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910026525835264
author Johnson, Bjorn
Levy, Jared
author_facet Johnson, Bjorn
Levy, Jared
contents Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a model based on current state of the art systems that fuses the two disciplines in order to perform effective speech to speech style transfer for the purpose of voice impersonation. We show that our proposed model is quite powerful, and succeeds in generating realistic audio samples despite a number of drawbacks in its capacity. We benchmark our proposed model by comparing it with a generative adversarial model which accomplishes a similar task, and show that ours produces more convincing results.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16721
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speech to Speech Synthesis for Voice Impersonation
Johnson, Bjorn
Levy, Jared
Sound
Machine Learning
Audio and Speech Processing
Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a model based on current state of the art systems that fuses the two disciplines in order to perform effective speech to speech style transfer for the purpose of voice impersonation. We show that our proposed model is quite powerful, and succeeds in generating realistic audio samples despite a number of drawbacks in its capacity. We benchmark our proposed model by comparing it with a generative adversarial model which accomplishes a similar task, and show that ours produces more convincing results.
title Speech to Speech Synthesis for Voice Impersonation
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2602.16721