Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mi, Qirui, Ma, Zhijian, Yang, Mengyue, Li, Haoxuan, Wang, Yisen, Zhang, Haifeng, Wang, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917543620378624
author Mi, Qirui
Ma, Zhijian
Yang, Mengyue
Li, Haoxuan
Wang, Yisen
Zhang, Haifeng
Wang, Jun
author_facet Mi, Qirui
Ma, Zhijian
Yang, Mengyue
Li, Haoxuan
Wang, Yisen
Zhang, Haifeng
Wang, Jun
contents LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. This insufficient experience reuse leads to computational redundancy and instability. To bridge this gap, we propose Skill-Pro, a framework enabling agents to autonomously learn reusable procedural skills from interaction experiences without parameter updates. By formalizing a Skill-MDP, Skill-Pro transforms passive episodic narratives into executable Skills defined by activation, execution, and termination conditions to ensure executability. To achieve reliable reusability without capability degradation, we introduce Non-Parametric PPO, which leverages semantic gradients for high-quality candidate generation and a PPO Gate for robust Skill verification. Through score-based maintenance, Skill-Pro sustains compact, high-quality procedural memory. Experimental results across in-domain, cross-task, and cross-agent scenarios demonstrate that Skill-Pro achieves superior reuse rates and significant gains with extreme memory compression. Visualized evolutionary trajectories and Skill distributions further reveal how Skill-Pro transparently accumulates, refines, and reuses procedural knowledge to facilitate long-term autonomy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01869
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
Mi, Qirui
Ma, Zhijian
Yang, Mengyue
Li, Haoxuan
Wang, Yisen
Zhang, Haifeng
Wang, Jun
Artificial Intelligence
LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. This insufficient experience reuse leads to computational redundancy and instability. To bridge this gap, we propose Skill-Pro, a framework enabling agents to autonomously learn reusable procedural skills from interaction experiences without parameter updates. By formalizing a Skill-MDP, Skill-Pro transforms passive episodic narratives into executable Skills defined by activation, execution, and termination conditions to ensure executability. To achieve reliable reusability without capability degradation, we introduce Non-Parametric PPO, which leverages semantic gradients for high-quality candidate generation and a PPO Gate for robust Skill verification. Through score-based maintenance, Skill-Pro sustains compact, high-quality procedural memory. Experimental results across in-domain, cross-task, and cross-agent scenarios demonstrate that Skill-Pro achieves superior reuse rates and significant gains with extreme memory compression. Visualized evolutionary trajectories and Skill distributions further reveal how Skill-Pro transparently accumulates, refines, and reuses procedural knowledge to facilitate long-term autonomy.
title Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2602.01869