Wings: Learning Multimodal LLMs without Text-only Forgetting

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yi-Kai, Lu, Shiyin, Li, Yang, Ma, Yanqing, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Zhan, De-Chuan, Ye, Han-Jia
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911906443296768
author Zhang, Yi-Kai
Lu, Shiyin
Li, Yang
Ma, Yanqing
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
author_facet Zhang, Yi-Kai
Lu, Shiyin
Li, Yang
Ma, Yanqing
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
contents Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM catastrophically forgets the text-only instructions, which do not include images and can be addressed within the initial LLM. In this paper, we present Wings, a novel MLLM that excels in both text-only dialogues and multimodal comprehension. Analyzing MLLM attention in multimodal instructions reveals that text-only forgetting is related to the attention shifts from pre-image to post-image text. From that, we construct extra modules that act as the boosted learner to compensate for the attention shift. The complementary visual and textual learners, like "wings" on either side, are connected in parallel within each layer's attention block. Initially, image and text inputs are aligned with visual learners operating alongside the main attention, balancing focus on visual elements. Textual learners are later collaboratively integrated with attention-based routing to blend the outputs of the visual and textual learners. We design the Low-Rank Residual Attention (LoRRA) to guarantee high efficiency for learners. Our experimental results demonstrate that Wings outperforms equally-scaled MLLMs in both text-only and visual question-answering tasks. On a newly constructed Interleaved Image-Text (IIT) benchmark, Wings exhibits superior performance from text-only-rich to multimodal-rich question-answering tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wings: Learning Multimodal LLMs without Text-only Forgetting
Zhang, Yi-Kai
Lu, Shiyin
Li, Yang
Ma, Yanqing
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
Computation and Language
Artificial Intelligence
Machine Learning
Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM catastrophically forgets the text-only instructions, which do not include images and can be addressed within the initial LLM. In this paper, we present Wings, a novel MLLM that excels in both text-only dialogues and multimodal comprehension. Analyzing MLLM attention in multimodal instructions reveals that text-only forgetting is related to the attention shifts from pre-image to post-image text. From that, we construct extra modules that act as the boosted learner to compensate for the attention shift. The complementary visual and textual learners, like "wings" on either side, are connected in parallel within each layer's attention block. Initially, image and text inputs are aligned with visual learners operating alongside the main attention, balancing focus on visual elements. Textual learners are later collaboratively integrated with attention-based routing to blend the outputs of the visual and textual learners. We design the Low-Rank Residual Attention (LoRRA) to guarantee high efficiency for learners. Our experimental results demonstrate that Wings outperforms equally-scaled MLLMs in both text-only and visual question-answering tasks. On a newly constructed Interleaved Image-Text (IIT) benchmark, Wings exhibits superior performance from text-only-rich to multimodal-rich question-answering tasks.
title Wings: Learning Multimodal LLMs without Text-only Forgetting
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.03496