DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhuohui, Cheng, Bin, He, Bin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911623253327872
author Zhang, Zhuohui
Cheng, Bin
He, Bin
author_facet Zhang, Zhuohui
Cheng, Bin
He, Bin
contents Building scalable and reusable multi-agent decision policies from offline datasets remains a challenge in offline multi-agent reinforcement learning (MARL), as existing methods often rely on fixed observation formats and action spaces that limit generalization. In contrast, large language models (LLMs) offer a flexible modeling interface that can naturally accommodate heterogeneous observations and actions. Motivated by this, we propose the Decision Language Model (DLM), which formulates multi-agent decision making as a dialogue-style sequence prediction problem under the centralized training with decentralized execution paradigm. DLM is trained in two stages: a supervised fine-tuning phase, which leverages dialogue-style datasets for centralized training with inter-agent context and generates executable actions from offline trajectories, followed by a group relative policy optimization phase to enhance robustness to out-of-distribution actions through lightweight reward functions. Experiments on multiple benchmarks show that a unified DLM outperforms strong offline MARL baselines and LLM-based conversational decision-making methods, while demonstrating strong zero-shot generalization to unseen scenarios across tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23557
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision Making
Zhang, Zhuohui
Cheng, Bin
He, Bin
Multiagent Systems
Artificial Intelligence
Building scalable and reusable multi-agent decision policies from offline datasets remains a challenge in offline multi-agent reinforcement learning (MARL), as existing methods often rely on fixed observation formats and action spaces that limit generalization. In contrast, large language models (LLMs) offer a flexible modeling interface that can naturally accommodate heterogeneous observations and actions. Motivated by this, we propose the Decision Language Model (DLM), which formulates multi-agent decision making as a dialogue-style sequence prediction problem under the centralized training with decentralized execution paradigm. DLM is trained in two stages: a supervised fine-tuning phase, which leverages dialogue-style datasets for centralized training with inter-agent context and generates executable actions from offline trajectories, followed by a group relative policy optimization phase to enhance robustness to out-of-distribution actions through lightweight reward functions. Experiments on multiple benchmarks show that a unified DLM outperforms strong offline MARL baselines and LLM-based conversational decision-making methods, while demonstrating strong zero-shot generalization to unseen scenarios across tasks.
title DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision Making
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2604.23557