LaMI: Augmenting Large Language Models via Late Multi-Image Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yariv, Guy, Schwartz, Idan, Adi, Yossi, Benaim, Sagie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918439138885632
author Yariv, Guy
Schwartz, Idan
Adi, Yossi
Benaim, Sagie
author_facet Yariv, Guy
Schwartz, Idan
Adi, Yossi
Benaim, Sagie
contents Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emperor penguin's belly?"). Visual Language Models (VLMs) perform better on visually grounded tasks but face two limitations: (i) often reduced performance on text-only commonsense reasoning compared to text-trained LLMs, and (ii) adapting newly released LLMs to vision input typically requires costly multimodal training. An alternative augments LLMs with test-time visual signals, improving visual commonsense without harming textual reasoning, but prior designs often rely on early fusion and a single image, which can be suboptimal. We propose a late multi-image fusion method: multiple images are generated from the text prompt with a lightweight parallel sampling, and their prediction probabilities are combined with those of a text-only LLM through a late-fusion layer that integrates projected visual features just before the final prediction. Across visual commonsense and NLP benchmarks, our method significantly outperforms augmented LLMs on visual reasoning, matches VLMs on vision-based tasks, and, when applied to strong LLMs such as LLaMA 3, also improves NLP performance while adding only modest test-time overhead. Project page is available at: https://guyyariv.github.io/LaMI.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
Yariv, Guy
Schwartz, Idan
Adi, Yossi
Benaim, Sagie
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emperor penguin's belly?"). Visual Language Models (VLMs) perform better on visually grounded tasks but face two limitations: (i) often reduced performance on text-only commonsense reasoning compared to text-trained LLMs, and (ii) adapting newly released LLMs to vision input typically requires costly multimodal training. An alternative augments LLMs with test-time visual signals, improving visual commonsense without harming textual reasoning, but prior designs often rely on early fusion and a single image, which can be suboptimal. We propose a late multi-image fusion method: multiple images are generated from the text prompt with a lightweight parallel sampling, and their prediction probabilities are combined with those of a text-only LLM through a late-fusion layer that integrates projected visual features just before the final prediction. Across visual commonsense and NLP benchmarks, our method significantly outperforms augmented LLMs on visual reasoning, matches VLMs on vision-based tasks, and, when applied to strong LLMs such as LLaMA 3, also improves NLP performance while adding only modest test-time overhead. Project page is available at: https://guyyariv.github.io/LaMI.
title LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2406.13621