What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Çoban, Enis Berk, Mandel, Michael I., Devaney, Johanna
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909219140141056
author Çoban, Enis Berk
Mandel, Michael I.
Devaney, Johanna
author_facet Çoban, Enis Berk
Mandel, Michael I.
Devaney, Johanna
contents Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound and images, known as multimodal LLMs (MLLMs), which are capable of describing images or sound recordings. Previous work has demonstrated that when the LLM component in MLLMs is frozen, the audio or visual encoder serves to caption the sound or image input facilitating text-based reasoning with the LLM component. We are interested in using the LLM's reasoning capabilities in order to facilitate classification. In this paper, we demonstrate through a captioning/classification experiment that an audio MLLM cannot fully leverage its LLM's text-based reasoning when generating audio captions. We also consider how this may be due to MLLMs separately representing auditory and textual information such that it severs the reasoning pathway from the LLM to the audio encoder.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models
Çoban, Enis Berk
Mandel, Michael I.
Devaney, Johanna
Audio and Speech Processing
Computation and Language
Sound
Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound and images, known as multimodal LLMs (MLLMs), which are capable of describing images or sound recordings. Previous work has demonstrated that when the LLM component in MLLMs is frozen, the audio or visual encoder serves to caption the sound or image input facilitating text-based reasoning with the LLM component. We are interested in using the LLM's reasoning capabilities in order to facilitate classification. In this paper, we demonstrate through a captioning/classification experiment that an audio MLLM cannot fully leverage its LLM's text-based reasoning when generating audio captions. We also consider how this may be due to MLLMs separately representing auditory and textual information such that it severs the reasoning pathway from the LLM to the audio encoder.
title What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2406.04615