Behavior Knowledge Merge in Reinforced Agentic Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yuan, Xiangchi, Shi, Dachuan, Zhang, Chunhui, Liu, Zheyuan, Yao, Shenglong, Vosoughi, Soroush, Lee, Wenke
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908775693156352
author Yuan, Xiangchi
Shi, Dachuan
Zhang, Chunhui
Liu, Zheyuan
Yao, Shenglong
Vosoughi, Soroush
Lee, Wenke
author_facet Yuan, Xiangchi
Shi, Dachuan
Zhang, Chunhui
Liu, Zheyuan
Yao, Shenglong
Vosoughi, Soroush
Lee, Wenke
contents Reinforcement learning (RL) is central to post-training, particularly for agentic models that require specialized reasoning behaviors. In this setting, model merging offers a practical mechanism for integrating multiple RL-trained agents from different tasks into a single generalist model. However, existing merging methods are designed for supervised fine-tuning (SFT), and they are suboptimal to preserve task-specific capabilities on RL-trained agentic models. The root is a task-vector mismatch between RL and SFT: on-policy RL induces task vectors that are highly sparse and heterogeneous, whereas SFT-style merging implicitly assumes dense and globally comparable task vectors. When standard global averaging is applied under this mismatch, RL's non-overlapping task vectors that encode critical task-specific behaviors are reduced and parameter updates are diluted. To address this issue, we propose Reinforced Agent Merging (RAM), a distribution-aware merging framework explicitly designed for RL-trained agentic models. RAM disentangles shared and task-specific unique parameter updates, averaging shared components while selectively preserving and rescaling unique ones to counteract parameter update dilution. Experiments across multiple agent domains and model architectures demonstrate that RAM not only surpasses merging baselines, but also unlocks synergistic potential among agents to achieve performance superior to that of specialized agents in their domains.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13572
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Behavior Knowledge Merge in Reinforced Agentic Models
Yuan, Xiangchi
Shi, Dachuan
Zhang, Chunhui
Liu, Zheyuan
Yao, Shenglong
Vosoughi, Soroush
Lee, Wenke
Machine Learning
Reinforcement learning (RL) is central to post-training, particularly for agentic models that require specialized reasoning behaviors. In this setting, model merging offers a practical mechanism for integrating multiple RL-trained agents from different tasks into a single generalist model. However, existing merging methods are designed for supervised fine-tuning (SFT), and they are suboptimal to preserve task-specific capabilities on RL-trained agentic models. The root is a task-vector mismatch between RL and SFT: on-policy RL induces task vectors that are highly sparse and heterogeneous, whereas SFT-style merging implicitly assumes dense and globally comparable task vectors. When standard global averaging is applied under this mismatch, RL's non-overlapping task vectors that encode critical task-specific behaviors are reduced and parameter updates are diluted. To address this issue, we propose Reinforced Agent Merging (RAM), a distribution-aware merging framework explicitly designed for RL-trained agentic models. RAM disentangles shared and task-specific unique parameter updates, averaging shared components while selectively preserving and rescaling unique ones to counteract parameter update dilution. Experiments across multiple agent domains and model architectures demonstrate that RAM not only surpasses merging baselines, but also unlocks synergistic potential among agents to achieve performance superior to that of specialized agents in their domains.
title Behavior Knowledge Merge in Reinforced Agentic Models
topic Machine Learning
url https://arxiv.org/abs/2601.13572