Saved in:
Bibliographic Details
Main Authors: Vedanshu, Tripathi, MM, Jaint, Bhavnesh
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.17813
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917732921901056
author Vedanshu
Tripathi, MM
Jaint, Bhavnesh
author_facet Vedanshu
Tripathi, MM
Jaint, Bhavnesh
contents The integration of large language models (LLMs) with vision-language (VL) tasks has been a transformative development in the realm of artificial intelligence, highlighting the potential of LLMs as a versatile general-purpose chatbot. However, the current trend in this evolution focuses on the integration of vision and language to create models that can operate in more diverse and real-world contexts. We present a novel approach, termed Bottleneck Adapter, specifically crafted for enhancing the multimodal functionalities of these complex models, enabling joint optimization of the entire multimodal LLM framework through a process known as Multimodal Model Tuning (MMT). Our approach utilizes lightweight adapters to connect the image encoder and LLM without the need for large, complex neural networks. Unlike the conventional modular training schemes, our approach adopts an end-to-end optimization regime, which, when combined with the adapters, facilitates the joint optimization using a significantly smaller parameter set. Our method exhibits robust performance with 90.12\% accuracy, outperforming both human-level performance (88.4\%) and LaVIN-7B (89.41\%).
format Preprint
id arxiv_https___arxiv_org_abs_2407_17813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning
Vedanshu
Tripathi, MM
Jaint, Bhavnesh
Computer Vision and Pattern Recognition
Artificial Intelligence
The integration of large language models (LLMs) with vision-language (VL) tasks has been a transformative development in the realm of artificial intelligence, highlighting the potential of LLMs as a versatile general-purpose chatbot. However, the current trend in this evolution focuses on the integration of vision and language to create models that can operate in more diverse and real-world contexts. We present a novel approach, termed Bottleneck Adapter, specifically crafted for enhancing the multimodal functionalities of these complex models, enabling joint optimization of the entire multimodal LLM framework through a process known as Multimodal Model Tuning (MMT). Our approach utilizes lightweight adapters to connect the image encoder and LLM without the need for large, complex neural networks. Unlike the conventional modular training schemes, our approach adopts an end-to-end optimization regime, which, when combined with the adapters, facilitates the joint optimization using a significantly smaller parameter set. Our method exhibits robust performance with 90.12\% accuracy, outperforming both human-level performance (88.4\%) and LaVIN-7B (89.41\%).
title Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.17813