Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Matsui, Yuta, Yamaki, Ryosuke, Ueda, Ryo, Shinagawa, Seitaro, Taniguchi, Tadahiro
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910910691409920
author Matsui, Yuta
Yamaki, Ryosuke
Ueda, Ryo
Shinagawa, Seitaro
Taniguchi, Tadahiro
author_facet Matsui, Yuta
Yamaki, Ryosuke
Ueda, Ryo
Shinagawa, Seitaro
Taniguchi, Tadahiro
contents We propose the Metropolis-Hastings Captioning Game (MHCG), a method to fuse knowledge of multiple vision-language models (VLMs) by learning from each other. Although existing methods that combine multiple models suffer from inference costs and architectural constraints, MHCG avoids these problems by performing decentralized Bayesian inference through a process resembling a language game. The knowledge fusion process establishes communication between two VLM agents alternately captioning images and learning from each other. We conduct two image-captioning experiments with two VLMs, each pre-trained on a different dataset. The first experiment demonstrates that MHCG achieves consistent improvement in reference-free evaluation metrics. The second experiment investigates how MHCG contributes to sharing VLMs' category-level vocabulary by observing the occurrence of the vocabulary in the generated captions.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09620
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
Matsui, Yuta
Yamaki, Ryosuke
Ueda, Ryo
Shinagawa, Seitaro
Taniguchi, Tadahiro
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multiagent Systems
We propose the Metropolis-Hastings Captioning Game (MHCG), a method to fuse knowledge of multiple vision-language models (VLMs) by learning from each other. Although existing methods that combine multiple models suffer from inference costs and architectural constraints, MHCG avoids these problems by performing decentralized Bayesian inference through a process resembling a language game. The knowledge fusion process establishes communication between two VLM agents alternately captioning images and learning from each other. We conduct two image-captioning experiments with two VLMs, each pre-trained on a different dataset. The first experiment demonstrates that MHCG achieves consistent improvement in reference-free evaluation metrics. The second experiment investigates how MHCG contributes to sharing VLMs' category-level vocabulary by observing the occurrence of the vocabulary in the generated captions.
title Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multiagent Systems
url https://arxiv.org/abs/2504.09620