Voice "Cloning" is Style Transfer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Kaitlyn, Bianchi, Federico, Bartelds, Martijn, Pot, Anna, Kwon, Yongchan, Zou, James
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911722474831872
author Zhou, Kaitlyn
Bianchi, Federico
Bartelds, Martijn
Pot, Anna
Kwon, Yongchan
Zou, James
author_facet Zhou, Kaitlyn
Bianchi, Federico
Bartelds, Martijn
Pot, Anna
Kwon, Yongchan
Zou, James
contents Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully ''clone'' an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16578
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Voice "Cloning" is Style Transfer
Zhou, Kaitlyn
Bianchi, Federico
Bartelds, Martijn
Pot, Anna
Kwon, Yongchan
Zou, James
Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully ''clone'' an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.
title Voice "Cloning" is Style Transfer
topic Sound
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2605.16578