RiTTA: Modeling Event Relations in Text-to-Audio Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yuhang, Jain, Yash, Liu, Xubo, Markham, Andrew, Vineet, Vibhav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913016912543744
author He, Yuhang
Jain, Yash
Liu, Xubo
Markham, Andrew
Vineet, Vibhav
author_facet He, Yuhang
Jain, Yash
Liu, Xubo
Markham, Andrew
Vineet, Vibhav
contents Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relations between audio events described in the input text. However, previous TTA methods have not systematically explored audio event relation modeling, nor have they proposed frameworks to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark for this task by: 1. proposing a comprehensive relation corpus covering all potential relations in real-world scenarios; 2. introducing a new audio event corpus encompassing commonly heard audios; and 3. proposing new evaluation metrics to assess audio event relation modeling from various perspectives. Furthermore, we propose a finetuning framework to enhance existing TTA models ability to model audio events relation. Code is available at: https://github.com/yuhanghe01/RiTTA
format Preprint
id arxiv_https___arxiv_org_abs_2412_15922
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RiTTA: Modeling Event Relations in Text-to-Audio Generation
He, Yuhang
Jain, Yash
Liu, Xubo
Markham, Andrew
Vineet, Vibhav
Machine Learning
Sound
Audio and Speech Processing
Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relations between audio events described in the input text. However, previous TTA methods have not systematically explored audio event relation modeling, nor have they proposed frameworks to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark for this task by: 1. proposing a comprehensive relation corpus covering all potential relations in real-world scenarios; 2. introducing a new audio event corpus encompassing commonly heard audios; and 3. proposing new evaluation metrics to assess audio event relation modeling from various perspectives. Furthermore, we propose a finetuning framework to enhance existing TTA models ability to model audio events relation. Code is available at: https://github.com/yuhanghe01/RiTTA
title RiTTA: Modeling Event Relations in Text-to-Audio Generation
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.15922