The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bai, Longju, Borah, Angana, Ignat, Oana, Mihalcea, Rada
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912123481751552
author Bai, Longju
Borah, Angana
Ignat, Oana
Mihalcea, Rada
author_facet Bai, Longju
Borah, Angana
Ignat, Oana
Mihalcea, Rada
contents Large Multimodal Models (LMMs) exhibit impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of most data and models. Conversely, multi-agent models have shown significant capability in solving complex tasks. Our study evaluates the collective performance of LMMs in a multi-agent interaction setting for the novel task of cultural image captioning. Our contributions are as follows: (1) We introduce MosAIC, a Multi-Agent framework to enhance cross-cultural Image Captioning using LMMs with distinct cultural personas; (2) We provide a dataset of culturally enriched image captions in English for images from China, India, and Romania across three datasets: GeoDE, GD-VCR, CVQA; (3) We propose a culture-adaptable metric for evaluating cultural information within image captions; and (4) We show that the multi-agent interaction outperforms single-agent models across different metrics, and offer valuable insights for future research. Our dataset and models can be accessed at https://github.com/MichiganNLP/MosAIC.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
Bai, Longju
Borah, Angana
Ignat, Oana
Mihalcea, Rada
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Large Multimodal Models (LMMs) exhibit impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of most data and models. Conversely, multi-agent models have shown significant capability in solving complex tasks. Our study evaluates the collective performance of LMMs in a multi-agent interaction setting for the novel task of cultural image captioning. Our contributions are as follows: (1) We introduce MosAIC, a Multi-Agent framework to enhance cross-cultural Image Captioning using LMMs with distinct cultural personas; (2) We provide a dataset of culturally enriched image captions in English for images from China, India, and Romania across three datasets: GeoDE, GD-VCR, CVQA; (3) We propose a culture-adaptable metric for evaluating cultural information within image captions; and (4) We show that the multi-agent interaction outperforms single-agent models across different metrics, and offer valuable insights for future research. Our dataset and models can be accessed at https://github.com/MichiganNLP/MosAIC.
title The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.11758