LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hinck, Musashi, Olson, Matthew L., Cobbley, David, Tseng, Shao-Yen, Lal, Vasudev
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917689710084096
author Hinck, Musashi
Olson, Matthew L.
Cobbley, David
Tseng, Shao-Yen
Lal, Vasudev
author_facet Hinck, Musashi
Olson, Matthew L.
Cobbley, David
Tseng, Shao-Yen
Lal, Vasudev
contents We train a suite of multimodal foundation models (MMFM) using the popular LLaVA framework with the recently released Gemma family of large language models (LLMs). Of particular interest is the 2B parameter Gemma model, which provides opportunities to construct capable small-scale MMFMs. In line with findings from other papers in this space, we test the effect of ablating three design features: pretraining the connector, utilizing a more powerful image backbone, and increasing the size of the language backbone. The resulting models, which we call LLaVA-Gemma, exhibit moderate performance on an array of evaluations, but fail to improve past the current comparably sized SOTA models. Closer analysis of performance shows mixed effects; skipping pretraining tends to reduce performance, larger vision models sometimes improve performance, and increasing language model size has inconsistent effects. We publicly release training recipes, code and weights for our models for the LLaVA-Gemma models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
Hinck, Musashi
Olson, Matthew L.
Cobbley, David
Tseng, Shao-Yen
Lal, Vasudev
Computation and Language
Artificial Intelligence
We train a suite of multimodal foundation models (MMFM) using the popular LLaVA framework with the recently released Gemma family of large language models (LLMs). Of particular interest is the 2B parameter Gemma model, which provides opportunities to construct capable small-scale MMFMs. In line with findings from other papers in this space, we test the effect of ablating three design features: pretraining the connector, utilizing a more powerful image backbone, and increasing the size of the language backbone. The resulting models, which we call LLaVA-Gemma, exhibit moderate performance on an array of evaluations, but fail to improve past the current comparably sized SOTA models. Closer analysis of performance shows mixed effects; skipping pretraining tends to reduce performance, larger vision models sometimes improve performance, and increasing language model size has inconsistent effects. We publicly release training recipes, code and weights for our models for the LLaVA-Gemma models.
title LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.01331