Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caren, Matthew, Chandra, Kartik, Tenenbaum, Joshua B., Ragan-Kelley, Jonathan, Ma, Karima
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910614513778688
author Caren, Matthew
Chandra, Kartik
Tenenbaum, Joshua B.
Ragan-Kelley, Jonathan
Ma, Karima
author_facet Caren, Matthew
Chandra, Kartik
Tenenbaum, Joshua B.
Ragan-Kelley, Jonathan
Ma, Karima
contents We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13507
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
Caren, Matthew
Chandra, Kartik
Tenenbaum, Joshua B.
Ragan-Kelley, Jonathan
Ma, Karima
Graphics
Computation and Language
Human-Computer Interaction
Sound
Audio and Speech Processing
I.3.8
We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.
title Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
topic Graphics
Computation and Language
Human-Computer Interaction
Sound
Audio and Speech Processing
I.3.8
url https://arxiv.org/abs/2409.13507