TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yi, Liu, Siao, Fan, Falong, Yao, Yuanqi, Zhao, Yue, Liu, Bo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914567799439360
author Xie, Yi
Liu, Siao
Fan, Falong
Yao, Yuanqi
Zhao, Yue
Liu, Bo
author_facet Xie, Yi
Liu, Siao
Fan, Falong
Yao, Yuanqi
Zhao, Yue
Liu, Bo
contents Multi-agent LLM systems have shown promise for complex reasoning, yet recent evaluations reveal they often underperform single-model baselines. We identify a structural failure mode in sequential fine-tuning of shared-context teams: updating one agent shifts the team's context distribution, and when subsequent updates are evaluated on cached rollouts, this mismatch compounds. We formalize this as the compounding occupancy shift and prove that stale-occupancy evaluation incurs a penalty that scales quadratically with the number of agents. In contrast, intermediate-occupancy evaluation reduces this to linear scaling. We propose TeamTR, a trust-region framework that resamples trajectories after each component update and enforces per-agent divergence control, yielding rigorous per-update and per-stage improvement lower bounds. Experiments show that TeamTR outperforms single-agent and sequential baselines with 7.1% on average, mitigates coordination regressions, and supports plug-and-play component replacement. Code is available at https://github.com/Yydc/TeamTR.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15207
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
Xie, Yi
Liu, Siao
Fan, Falong
Yao, Yuanqi
Zhao, Yue
Liu, Bo
Machine Learning
Multiagent Systems
Multi-agent LLM systems have shown promise for complex reasoning, yet recent evaluations reveal they often underperform single-model baselines. We identify a structural failure mode in sequential fine-tuning of shared-context teams: updating one agent shifts the team's context distribution, and when subsequent updates are evaluated on cached rollouts, this mismatch compounds. We formalize this as the compounding occupancy shift and prove that stale-occupancy evaluation incurs a penalty that scales quadratically with the number of agents. In contrast, intermediate-occupancy evaluation reduces this to linear scaling. We propose TeamTR, a trust-region framework that resamples trajectories after each component update and enforces per-agent divergence control, yielding rigorous per-update and per-stage improvement lower bounds. Experiments show that TeamTR outperforms single-agent and sequential baselines with 7.1% on average, mitigates coordination regressions, and supports plug-and-play component replacement. Code is available at https://github.com/Yydc/TeamTR.
title TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2605.15207