Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Eric Hanchen, Li, Levina, Sun, Rui, Liang, Xiao, Li, Yubei, Wu, Yuchen, Luo, Haozheng, Li, Hengli, Zhang, Zhi, Kang, Zhaolu, Chang, Kai-Wei, Wu, Ying Nian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908928192806912
author Jiang, Eric Hanchen
Li, Levina
Sun, Rui
Liang, Xiao
Li, Yubei
Wu, Yuchen
Luo, Haozheng
Li, Hengli
Zhang, Zhi
Kang, Zhaolu
Chang, Kai-Wei
Wu, Ying Nian
author_facet Jiang, Eric Hanchen
Li, Levina
Sun, Rui
Liang, Xiao
Li, Yubei
Wu, Yuchen
Luo, Haozheng
Li, Hengli
Zhang, Zhi
Kang, Zhaolu
Chang, Kai-Wei
Wu, Ying Nian
contents Large Language Models (LLMs) have shown remarkable performance in completing various tasks. However, solving complex problems often requires the coordination of multiple agents, raising a fundamental question: how to effectively select and interconnect these agents. In this paper, we propose \textbf{Agent Q-Mix}, a reinforcement learning framework that reformulates topology selection as a cooperative Multi-Agent Reinforcement Learning (MARL) problem. Our method learns decentralized communication decisions using QMIX value factorization, where each agent selects from a set of communication actions that jointly induce a round-wise communication graph. At its core, Agent Q-Mix combines a topology-aware GNN encoder, GRU memory, and per-agent Q-heads under a Centralized Training with Decentralized Execution (CTDE) paradigm. The framework optimizes a reward function that balances task accuracy with token cost. Across seven core benchmarks in coding, reasoning, and mathematics, Agent Q-Mix achieves the highest average accuracy compared to existing methods while demonstrating superior token efficiency and robustness against agent failure. Notably, on the challenging Humanity's Last Exam (HLE) using Gemini-3.1-Flash-Lite as a backbone, Agent Q-Mix achieves 20.8\% accuracy, outperforming Microsoft Agent Framework (19.2\%) and LangGraph (19.2\%), followed by AutoGen and Lobster by OpenClaw. These results underscore the effectiveness of learned, decentralized topology optimization in pushing the boundaries of multi-agent reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00344
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning
Jiang, Eric Hanchen
Li, Levina
Sun, Rui
Liang, Xiao
Li, Yubei
Wu, Yuchen
Luo, Haozheng
Li, Hengli
Zhang, Zhi
Kang, Zhaolu
Chang, Kai-Wei
Wu, Ying Nian
Computation and Language
Applications
Large Language Models (LLMs) have shown remarkable performance in completing various tasks. However, solving complex problems often requires the coordination of multiple agents, raising a fundamental question: how to effectively select and interconnect these agents. In this paper, we propose \textbf{Agent Q-Mix}, a reinforcement learning framework that reformulates topology selection as a cooperative Multi-Agent Reinforcement Learning (MARL) problem. Our method learns decentralized communication decisions using QMIX value factorization, where each agent selects from a set of communication actions that jointly induce a round-wise communication graph. At its core, Agent Q-Mix combines a topology-aware GNN encoder, GRU memory, and per-agent Q-heads under a Centralized Training with Decentralized Execution (CTDE) paradigm. The framework optimizes a reward function that balances task accuracy with token cost. Across seven core benchmarks in coding, reasoning, and mathematics, Agent Q-Mix achieves the highest average accuracy compared to existing methods while demonstrating superior token efficiency and robustness against agent failure. Notably, on the challenging Humanity's Last Exam (HLE) using Gemini-3.1-Flash-Lite as a backbone, Agent Q-Mix achieves 20.8\% accuracy, outperforming Microsoft Agent Framework (19.2\%) and LangGraph (19.2\%), followed by AutoGen and Lobster by OpenClaw. These results underscore the effectiveness of learned, decentralized topology optimization in pushing the boundaries of multi-agent reasoning.
title Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning
topic Computation and Language
Applications
url https://arxiv.org/abs/2604.00344