An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiawei, Çoban, Enis Berk, Schevchenko, Zarina, Tang, Hao, Zhu, Zhigang, Mandel, Michael I, Devaney, Johanna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911248042426368
author Liu, Jiawei
Çoban, Enis Berk
Schevchenko, Zarina
Tang, Hao
Zhu, Zhigang
Mandel, Michael I
Devaney, Johanna
author_facet Liu, Jiawei
Çoban, Enis Berk
Schevchenko, Zarina
Tang, Hao
Zhu, Zhigang
Mandel, Michael I
Devaney, Johanna
contents Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not encourage deep integration of modalities, limiting the model's ability to leverage the core language model's reasoning capabilities. This work examined the impact of interleaved instruction tuning in an audio MLLM, where audio tokens are interleaved within the prompt. Using the Listen, Think, and Understand (LTU) model as a testbed, we conduct an experiment using the Synonym and Hypernym Audio Reasoning Dataset (SHARD), our newly created reasoning benchmark for audio-based semantic reasoning focusing on synonym and hypernym recognition. Our findings show that while even zero-shot interleaved prompting improves performance on our reasoning tasks, a small amount of fine-tuning using interleaved training prompts improves the results further, however, at the expense of the MLLM's audio labeling ability.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
Liu, Jiawei
Çoban, Enis Berk
Schevchenko, Zarina
Tang, Hao
Zhu, Zhigang
Mandel, Michael I
Devaney, Johanna
Multimedia
Computation and Language
Sound
Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not encourage deep integration of modalities, limiting the model's ability to leverage the core language model's reasoning capabilities. This work examined the impact of interleaved instruction tuning in an audio MLLM, where audio tokens are interleaved within the prompt. Using the Listen, Think, and Understand (LTU) model as a testbed, we conduct an experiment using the Synonym and Hypernym Audio Reasoning Dataset (SHARD), our newly created reasoning benchmark for audio-based semantic reasoning focusing on synonym and hypernym recognition. Our findings show that while even zero-shot interleaved prompting improves performance on our reasoning tasks, a small amount of fine-tuning using interleaved training prompts improves the results further, however, at the expense of the MLLM's audio labeling ability.
title An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
topic Multimedia
Computation and Language
Sound
url https://arxiv.org/abs/2511.02234