Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jiang, Bowen, Zhuang, Zhijun, Shivakumar, Shreyas S., Roth, Dan, Taylor, Camillo J.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916171141349376
author Jiang, Bowen
Zhuang, Zhijun
Shivakumar, Shreyas S.
Roth, Dan
Taylor, Camillo J.
author_facet Jiang, Bowen
Zhuang, Zhijun
Shivakumar, Shreyas S.
Roth, Dan
Taylor, Camillo J.
contents This work explores the zero-shot capabilities of foundation models in Visual Question Answering (VQA) tasks. We propose an adaptive multi-agent system, named Multi-Agent VQA, to overcome the limitations of foundation models in object detection and counting by using specialized agents as tools. Unlike existing approaches, our study focuses on the system's performance without fine-tuning it on specific VQA datasets, making it more practical and robust in the open world. We present preliminary experimental results under zero-shot scenarios and highlight some failure cases, offering new directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14783
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
Jiang, Bowen
Zhuang, Zhijun
Shivakumar, Shreyas S.
Roth, Dan
Taylor, Camillo J.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multiagent Systems
This work explores the zero-shot capabilities of foundation models in Visual Question Answering (VQA) tasks. We propose an adaptive multi-agent system, named Multi-Agent VQA, to overcome the limitations of foundation models in object detection and counting by using specialized agents as tools. Unlike existing approaches, our study focuses on the system's performance without fine-tuning it on specific VQA datasets, making it more practical and robust in the open world. We present preliminary experimental results under zero-shot scenarios and highlight some failure cases, offering new directions for future research.
title Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2403.14783