UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yiqun, Yang, Wei, Zhang, Erhan, Wang, Shijie, Liu, Qi, Niu, Zechun, Zhang, Bin, Li, Haitao, Li, Rui, Yan, Lingyong, Feng, Jinyuan, Qi, Biqing, Wei, Xiaochi, Gao, Yan, Wu, Yi, Hu, Yao, Mao, Jiaxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910258329288704
author Chen, Yiqun
Yang, Wei
Zhang, Erhan
Wang, Shijie
Liu, Qi
Niu, Zechun
Zhang, Bin
Li, Haitao
Li, Rui
Yan, Lingyong
Feng, Jinyuan
Qi, Biqing
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Mao, Jiaxin
author_facet Chen, Yiqun
Yang, Wei
Zhang, Erhan
Wang, Shijie
Liu, Qi
Niu, Zechun
Zhang, Bin
Li, Haitao
Li, Rui
Yan, Lingyong
Feng, Jinyuan
Qi, Biqing
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Mao, Jiaxin
contents LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules, while agents are rarely optimized through a unified reinforcement learning interface. Existing RL post-training frameworks mainly target single-policy optimization and lack abstractions for user-defined multi-agent workflows, structured interaction, role-specific credit assignment, and configurable parameter sharing. We present UnityMAS-O, a general RL optimization framework for LLM-based multi-agent systems. UnityMAS-O treats the complete workflow as the optimization unit, rather than a single response or policy trajectory. It represents workflows through four first-class objects: logical agent roles, graph trajectories, user-defined rewards, and agent--model mappings. This decouples logical agents from physical model parameters, supporting full sharing, full separation, and partial sharing, with rewards assigned at role, turn, and trajectory levels. UnityMAS-O extends verl with a Ray-based star-topology runtime. A central controller executes workflows, invokes tools, records structured trajectories, and assembles rewards; model-local worker groups handle rollout, buffering, advantage computation, and distributed PPO-style updates. Users can define agents, workflows, model mappings, and rewards without rewriting the optimization infrastructure. We instantiate UnityMAS-O on retrieval-augmented QA, iterative agentic search, and reflective code generation. Across Natural Questions, HotpotQA, and held-out code tasks, multi-agent RL improves manually specified workflows after optimization, with especially large gains for smaller models and strict code all-passed metrics. These results show that UnityMAS-O can serve as a reusable substrate for converting diverse LLM-based multi-agent workflows into trainable multi-agent RL systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26646
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
Chen, Yiqun
Yang, Wei
Zhang, Erhan
Wang, Shijie
Liu, Qi
Niu, Zechun
Zhang, Bin
Li, Haitao
Li, Rui
Yan, Lingyong
Feng, Jinyuan
Qi, Biqing
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Mao, Jiaxin
Artificial Intelligence
Computation and Language
Multiagent Systems
LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules, while agents are rarely optimized through a unified reinforcement learning interface. Existing RL post-training frameworks mainly target single-policy optimization and lack abstractions for user-defined multi-agent workflows, structured interaction, role-specific credit assignment, and configurable parameter sharing. We present UnityMAS-O, a general RL optimization framework for LLM-based multi-agent systems. UnityMAS-O treats the complete workflow as the optimization unit, rather than a single response or policy trajectory. It represents workflows through four first-class objects: logical agent roles, graph trajectories, user-defined rewards, and agent--model mappings. This decouples logical agents from physical model parameters, supporting full sharing, full separation, and partial sharing, with rewards assigned at role, turn, and trajectory levels. UnityMAS-O extends verl with a Ray-based star-topology runtime. A central controller executes workflows, invokes tools, records structured trajectories, and assembles rewards; model-local worker groups handle rollout, buffering, advantage computation, and distributed PPO-style updates. Users can define agents, workflows, model mappings, and rewards without rewriting the optimization infrastructure. We instantiate UnityMAS-O on retrieval-augmented QA, iterative agentic search, and reflective code generation. Across Natural Questions, HotpotQA, and held-out code tasks, multi-agent RL improves manually specified workflows after optimization, with especially large gains for smaller models and strict code all-passed metrics. These results show that UnityMAS-O can serve as a reusable substrate for converting diverse LLM-based multi-agent workflows into trainable multi-agent RL systems.
title UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2605.26646