Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912202653433856 |
|---|---|
| author | Mahmud, Saaduddin Goldfajn, Dorian Benhamou Zilberstein, Shlomo |
| author_facet | Mahmud, Saaduddin Goldfajn, Dorian Benhamou Zilberstein, Shlomo |
| contents | Distributed Constraint Optimization Problems (DCOPs) offer a powerful framework for multi-agent coordination but often rely on labor-intensive, manual problem construction. To address this, we introduce VL-DCOPs, a framework that takes advantage of large multimodal foundation models (LFMs) to automatically generate constraints from both visual and linguistic instructions. We then introduce a spectrum of agent archetypes for solving VL-DCOPs: from a neuro-symbolic agent that delegates some of the algorithmic decisions to an LFM, to a fully neural agent that depends entirely on an LFM for coordination. We evaluate these agent archetypes using state-of-the-art LLMs (large language models) and VLMs (vision language models) on three novel VL-DCOP tasks and compare their respective advantages and drawbacks. Lastly, we discuss how this work extends to broader frontier challenges in the DCOP literature. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_14189 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models Mahmud, Saaduddin Goldfajn, Dorian Benhamou Zilberstein, Shlomo Artificial Intelligence Machine Learning Multiagent Systems Distributed Constraint Optimization Problems (DCOPs) offer a powerful framework for multi-agent coordination but often rely on labor-intensive, manual problem construction. To address this, we introduce VL-DCOPs, a framework that takes advantage of large multimodal foundation models (LFMs) to automatically generate constraints from both visual and linguistic instructions. We then introduce a spectrum of agent archetypes for solving VL-DCOPs: from a neuro-symbolic agent that delegates some of the algorithmic decisions to an LFM, to a fully neural agent that depends entirely on an LFM for coordination. We evaluate these agent archetypes using state-of-the-art LLMs (large language models) and VLMs (vision language models) on three novel VL-DCOP tasks and compare their respective advantages and drawbacks. Lastly, we discuss how this work extends to broader frontier challenges in the DCOP literature. |
| title | Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models |
| topic | Artificial Intelligence Machine Learning Multiagent Systems |
| url | https://arxiv.org/abs/2501.14189 |