Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahmud, Saaduddin, Goldfajn, Dorian Benhamou, Zilberstein, Shlomo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912202653433856
author Mahmud, Saaduddin
Goldfajn, Dorian Benhamou
Zilberstein, Shlomo
author_facet Mahmud, Saaduddin
Goldfajn, Dorian Benhamou
Zilberstein, Shlomo
contents Distributed Constraint Optimization Problems (DCOPs) offer a powerful framework for multi-agent coordination but often rely on labor-intensive, manual problem construction. To address this, we introduce VL-DCOPs, a framework that takes advantage of large multimodal foundation models (LFMs) to automatically generate constraints from both visual and linguistic instructions. We then introduce a spectrum of agent archetypes for solving VL-DCOPs: from a neuro-symbolic agent that delegates some of the algorithmic decisions to an LFM, to a fully neural agent that depends entirely on an LFM for coordination. We evaluate these agent archetypes using state-of-the-art LLMs (large language models) and VLMs (vision language models) on three novel VL-DCOP tasks and compare their respective advantages and drawbacks. Lastly, we discuss how this work extends to broader frontier challenges in the DCOP literature.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models
Mahmud, Saaduddin
Goldfajn, Dorian Benhamou
Zilberstein, Shlomo
Artificial Intelligence
Machine Learning
Multiagent Systems
Distributed Constraint Optimization Problems (DCOPs) offer a powerful framework for multi-agent coordination but often rely on labor-intensive, manual problem construction. To address this, we introduce VL-DCOPs, a framework that takes advantage of large multimodal foundation models (LFMs) to automatically generate constraints from both visual and linguistic instructions. We then introduce a spectrum of agent archetypes for solving VL-DCOPs: from a neuro-symbolic agent that delegates some of the algorithmic decisions to an LFM, to a fully neural agent that depends entirely on an LFM for coordination. We evaluate these agent archetypes using state-of-the-art LLMs (large language models) and VLMs (vision language models) on three novel VL-DCOP tasks and compare their respective advantages and drawbacks. Lastly, we discuss how this work extends to broader frontier challenges in the DCOP literature.
title Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models
topic Artificial Intelligence
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2501.14189