When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Minsu, Lee, Seolhee, Lee, Choonghyeon, Cho, Namhyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910975258525696
author Kang, Minsu
Lee, Seolhee
Lee, Choonghyeon
Cho, Namhyun
author_facet Kang, Minsu
Lee, Seolhee
Lee, Choonghyeon
Cho, Namhyun
contents Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broader range of non-speech sounds, including natural sounds (lion-roars, birdsongs) and designed voice (synthetic growls). To accomodate generation of diverse non-speech sounds and 44.1kHz high-quality audio transformation, we introduce a preprocessing pipeline and an improved CVAE-based H2NH-VC model, both optimized for human and non-human voices. Experimental results showed that the proposed method outperformed baselines in quality, naturalness, and similarity MOS, achieving effective voice conversion across diverse non-human timbres. Demo samples are available at https://nc-ai.github.io/speech/publications/nonhuman-vc/
format Preprint
id arxiv_https___arxiv_org_abs_2505_24336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
Kang, Minsu
Lee, Seolhee
Lee, Choonghyeon
Cho, Namhyun
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Signal Processing
Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broader range of non-speech sounds, including natural sounds (lion-roars, birdsongs) and designed voice (synthetic growls). To accomodate generation of diverse non-speech sounds and 44.1kHz high-quality audio transformation, we introduce a preprocessing pipeline and an improved CVAE-based H2NH-VC model, both optimized for human and non-human voices. Experimental results showed that the proposed method outperformed baselines in quality, naturalness, and similarity MOS, achieving effective voice conversion across diverse non-human timbres. Demo samples are available at https://nc-ai.github.io/speech/publications/nonhuman-vc/
title When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2505.24336