Example-Based Framework for Perceptually Guided Audio Texture Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kamath, Purnima, Gupta, Chitralekha, Wyse, Lonce, Nanayakkara, Suranga
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914964860567552
author Kamath, Purnima
Gupta, Chitralekha
Wyse, Lonce
Nanayakkara, Suranga
author_facet Kamath, Purnima
Gupta, Chitralekha
Wyse, Lonce
Nanayakkara, Suranga
contents Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we develop a method for semantic control over an unconditionally trained StyleGAN in the absence of such labeled datasets. In this paper, we propose an example-based framework to determine guidance vectors for audio texture generation based on user-defined semantic attributes. Our approach leverages the semantically disentangled latent space of an unconditionally trained StyleGAN. By using a few synthetic examples to indicate the presence or absence of a semantic attribute, we infer the guidance vectors in the latent space of the StyleGAN to control that attribute during generation. Our results show that our framework can find user-defined and perceptually relevant guidance vectors for controllable generation for audio textures. Furthermore, we demonstrate an application of our framework to other tasks, such as selective semantic attribute transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2308_11859
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Example-Based Framework for Perceptually Guided Audio Texture Generation
Kamath, Purnima
Gupta, Chitralekha
Wyse, Lonce
Nanayakkara, Suranga
Audio and Speech Processing
Artificial Intelligence
Sound
Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we develop a method for semantic control over an unconditionally trained StyleGAN in the absence of such labeled datasets. In this paper, we propose an example-based framework to determine guidance vectors for audio texture generation based on user-defined semantic attributes. Our approach leverages the semantically disentangled latent space of an unconditionally trained StyleGAN. By using a few synthetic examples to indicate the presence or absence of a semantic attribute, we infer the guidance vectors in the latent space of the StyleGAN to control that attribute during generation. Our results show that our framework can find user-defined and perceptually relevant guidance vectors for controllable generation for audio textures. Furthermore, we demonstrate an application of our framework to other tasks, such as selective semantic attribute transfer.
title Example-Based Framework for Perceptually Guided Audio Texture Generation
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2308.11859