Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kong, Mingze, Qu, Zikun, Zhou, Zhongquan, Liang, Pengyu, Li, Xiang, Shang, Zhiwei, Hong, Zhi, Huang, Kaiyu, Wang, Zhiyong, Dai, Zhongxiang
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915767309565952
author Kong, Mingze
Qu, Zikun
Zhou, Zhongquan
Liang, Pengyu
Li, Xiang
Shang, Zhiwei
Hong, Zhi
Huang, Kaiyu
Wang, Zhiyong
Dai, Zhongxiang
author_facet Kong, Mingze
Qu, Zikun
Zhou, Zhongquan
Liang, Pengyu
Li, Xiang
Shang, Zhiwei
Hong, Zhi
Huang, Kaiyu
Wang, Zhiyong
Dai, Zhongxiang
contents The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization methods typically formulate workflow synthesis as a static, one-shot code-centric generation problem. This paradigm imposes excessive constraints on the model's coding capabilities and restricts the flexibility required for dynamic problem-solving. In this paper, we present Workflow-R1, a framework that reformulates workflow construction as a multi-turn, natural language-based sequential decision-making process. To resolve the optimization granularity mismatch inherent in such multi-turn interactions, we introduce Group Sub-sequence Policy Optimization (GSsPO). While explicitly tailored to align with the interleaved Think-Action dynamics of agentic reasoning, GSsPO fundamentally functions as a structure-aware RL algorithm generalizable to a broad class of multi-turn agentic sequential decision-making tasks. By recalibrating the optimization unit to the composite sub-sequence, specifically the atomic Think-Action cycle, it aligns gradient updates with the semantic boundaries of these interactions, ensuring robust learning in complex multi-turn reasoning tasks. Through extensive experiments on multiple QA benchmarks, Workflow-R1 outperforms competitive baselines, validating GSsPO as a generalized solution for sequential reasoning and establishing Workflow-R1 as a promising new paradigm for automated workflow optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01202
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
Kong, Mingze
Qu, Zikun
Zhou, Zhongquan
Liang, Pengyu
Li, Xiang
Shang, Zhiwei
Hong, Zhi
Huang, Kaiyu
Wang, Zhiyong
Dai, Zhongxiang
Artificial Intelligence
The rapid evolution of agentic workflows has demonstrated strong performance of LLM-based agents in addressing complex reasoning tasks. However, existing workflow optimization methods typically formulate workflow synthesis as a static, one-shot code-centric generation problem. This paradigm imposes excessive constraints on the model's coding capabilities and restricts the flexibility required for dynamic problem-solving. In this paper, we present Workflow-R1, a framework that reformulates workflow construction as a multi-turn, natural language-based sequential decision-making process. To resolve the optimization granularity mismatch inherent in such multi-turn interactions, we introduce Group Sub-sequence Policy Optimization (GSsPO). While explicitly tailored to align with the interleaved Think-Action dynamics of agentic reasoning, GSsPO fundamentally functions as a structure-aware RL algorithm generalizable to a broad class of multi-turn agentic sequential decision-making tasks. By recalibrating the optimization unit to the composite sub-sequence, specifically the atomic Think-Action cycle, it aligns gradient updates with the semantic boundaries of these interactions, ensuring robust learning in complex multi-turn reasoning tasks. Through extensive experiments on multiple QA benchmarks, Workflow-R1 outperforms competitive baselines, validating GSsPO as a generalized solution for sequential reasoning and establishing Workflow-R1 as a promising new paradigm for automated workflow optimization.
title Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction
topic Artificial Intelligence
url https://arxiv.org/abs/2602.01202