Dissecting Temporal Understanding in Text-to-Audio Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oncescu, Andreea-Maria, Henriques, João F., Koepke, A. Sophia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916377331236864
author Oncescu, Andreea-Maria
Henriques, João F.
Koepke, A. Sophia
author_facet Oncescu, Andreea-Maria
Henriques, João F.
Koepke, A. Sophia
contents Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data, including objects, and characters. The models also need to learn spatial arrangements and temporal relationships. In this work, we analyse the temporal ordering of sounds, which is an understudied problem in the context of text-to-audio retrieval. In particular, we dissect the temporal understanding capabilities of a state-of-the-art model for text-to-audio retrieval on the AudioCaps and Clotho datasets. Additionally, we introduce a synthetic text-audio dataset that provides a controlled setting for evaluating temporal capabilities of recent models. Lastly, we present a loss function that encourages text-audio models to focus on the temporal ordering of events. Code and data are available at https://www.robots.ox.ac.uk/~vgg/research/audio-retrieval/dtu/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dissecting Temporal Understanding in Text-to-Audio Retrieval
Oncescu, Andreea-Maria
Henriques, João F.
Koepke, A. Sophia
Information Retrieval
Machine Learning
Sound
Audio and Speech Processing
Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data, including objects, and characters. The models also need to learn spatial arrangements and temporal relationships. In this work, we analyse the temporal ordering of sounds, which is an understudied problem in the context of text-to-audio retrieval. In particular, we dissect the temporal understanding capabilities of a state-of-the-art model for text-to-audio retrieval on the AudioCaps and Clotho datasets. Additionally, we introduce a synthetic text-audio dataset that provides a controlled setting for evaluating temporal capabilities of recent models. Lastly, we present a loss function that encourages text-audio models to focus on the temporal ordering of events. Code and data are available at https://www.robots.ox.ac.uk/~vgg/research/audio-retrieval/dtu/.
title Dissecting Temporal Understanding in Text-to-Audio Retrieval
topic Information Retrieval
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.00851