Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shitian, Zhang, Renrui, Luo, Xu, Wang, Yan, Zhang, Shanghang, Gao, Peng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914961316380672
author Zhao, Shitian
Zhang, Renrui
Luo, Xu
Wang, Yan
Zhang, Shanghang
Gao, Peng
author_facet Zhao, Shitian
Zhang, Renrui
Luo, Xu
Wang, Yan
Zhang, Shanghang
Gao, Peng
contents Model fusing has always been an important topic, especially in an era where large language models (LLM) and multi-modal language models (MLM) with different architectures, parameter sizes and training pipelines, are being created all the time. In this work, we propose a post-hoc framework, aiming at fusing heterogeneous models off-the-shell, which we call \textit{likelihood composition}, and the basic idea is to compose multiple models' likelihood distribution when doing a multi-choice visual-question-answering task. Here the core concept, \textit{likelihood}, is actually the log-probability of the candidate answer. In \textit{likelihood composition}, we introduce some basic operations: \textit{debias}, \textit{highlight}, \textit{majority-vote} and \textit{ensemble}. By combining (composing) these basic elements, we get the mixed composition methods: \textit{mix-composition}. Through conducting comprehensive experiments on 9 VQA datasets and 10 MLMs, we prove the effectiveness of \textit{mix-composition} compared with simple \textit{ensemble} or \textit{majority-vote} methods. In this framework, people can propose new basic composition methods and combine them to get the new mixed composition methods. We hope our proposed \textit{likelihood composition} can provide a new perspective of fusing heterogeneous models and inspire the exploration under this framework.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
Zhao, Shitian
Zhang, Renrui
Luo, Xu
Wang, Yan
Zhang, Shanghang
Gao, Peng
Computation and Language
Model fusing has always been an important topic, especially in an era where large language models (LLM) and multi-modal language models (MLM) with different architectures, parameter sizes and training pipelines, are being created all the time. In this work, we propose a post-hoc framework, aiming at fusing heterogeneous models off-the-shell, which we call \textit{likelihood composition}, and the basic idea is to compose multiple models' likelihood distribution when doing a multi-choice visual-question-answering task. Here the core concept, \textit{likelihood}, is actually the log-probability of the candidate answer. In \textit{likelihood composition}, we introduce some basic operations: \textit{debias}, \textit{highlight}, \textit{majority-vote} and \textit{ensemble}. By combining (composing) these basic elements, we get the mixed composition methods: \textit{mix-composition}. Through conducting comprehensive experiments on 9 VQA datasets and 10 MLMs, we prove the effectiveness of \textit{mix-composition} compared with simple \textit{ensemble} or \textit{majority-vote} methods. In this framework, people can propose new basic composition methods and combine them to get the new mixed composition methods. We hope our proposed \textit{likelihood composition} can provide a new perspective of fusing heterogeneous models and inspire the exploration under this framework.
title Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
topic Computation and Language
url https://arxiv.org/abs/2410.00363