Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chowdhury, Sanjoy, Nag, Sayan, Dasgupta, Subhrajyoti, Chen, Jun, Elhoseiny, Mohamed, Gao, Ruohan, Manocha, Dinesh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914857143500800
author Chowdhury, Sanjoy
Nag, Sayan
Dasgupta, Subhrajyoti
Chen, Jun
Elhoseiny, Mohamed
Gao, Ruohan
Manocha, Dinesh
author_facet Chowdhury, Sanjoy
Nag, Sayan
Dasgupta, Subhrajyoti
Chen, Jun
Elhoseiny, Mohamed
Gao, Ruohan
Manocha, Dinesh
contents Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding of the audio-visual semantics. We present Meerkat, an audio-visual LLM equipped with a fine-grained understanding of image and audio both spatially and temporally. With a new modality alignment module based on optimal transport and a cross-attention module that enforces audio-visual consistency, Meerkat can tackle challenging tasks such as audio referred image grounding, image guided audio temporal localization, and audio-visual fact-checking. Moreover, we carefully curate a large dataset AVFIT that comprises 3M instruction tuning samples collected from open-source datasets, and introduce MeerkatBench that unifies five challenging audio-visual tasks. We achieve state-of-the-art performance on all these downstream tasks with a relative improvement of up to 37.12%.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
Chowdhury, Sanjoy
Nag, Sayan
Dasgupta, Subhrajyoti
Chen, Jun
Elhoseiny, Mohamed
Gao, Ruohan
Manocha, Dinesh
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding of the audio-visual semantics. We present Meerkat, an audio-visual LLM equipped with a fine-grained understanding of image and audio both spatially and temporally. With a new modality alignment module based on optimal transport and a cross-attention module that enforces audio-visual consistency, Meerkat can tackle challenging tasks such as audio referred image grounding, image guided audio temporal localization, and audio-visual fact-checking. Moreover, we carefully curate a large dataset AVFIT that comprises 3M instruction tuning samples collected from open-source datasets, and introduce MeerkatBench that unifies five challenging audio-visual tasks. We achieve state-of-the-art performance on all these downstream tasks with a relative improvement of up to 37.12%.
title Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2407.01851