SkillOS: Learning Skill Curation for Self-Evolving Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ouyang, Siru, Yan, Jun, Chen, Yanfei, Han, Rujun, Wang, Zifeng, Mishra, Bhavana Dalvi, Meng, Rui, Li, Chun-Liang, Jiao, Yizhu, Zha, Kaiwen, Shen, Maohao, Tirumalashetty, Vishy, Lee, George, Han, Jiawei, Pfister, Tomas, Lee, Chen-Yu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910198949478400
author Ouyang, Siru
Yan, Jun
Chen, Yanfei
Han, Rujun
Wang, Zifeng
Mishra, Bhavana Dalvi
Meng, Rui
Li, Chun-Liang
Jiao, Yizhu
Zha, Kaiwen
Shen, Maohao
Tirumalashetty, Vishy
Lee, George
Han, Jiawei
Pfister, Tomas
Lee, Chen-Yu
author_facet Ouyang, Siru
Yan, Jun
Chen, Yanfei
Han, Rujun
Wang, Zifeng
Mishra, Bhavana Dalvi
Meng, Rui
Li, Chun-Liang
Jiao, Yizhu
Zha, Kaiwen
Shen, Maohao
Tirumalashetty, Vishy
Lee, George
Han, Jiawei
Pfister, Tomas
Lee, Chen-Yu
contents LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate for self-evolution, where high-quality skill curation serves as the key bottleneck. Existing approaches either rely on manual skill curation, prescribe heuristic skill operations, or train for short-horizon skill operations. However, they still struggle to learn complex long-term curation policies from indirect and delayed feedback. To tackle this challenge, we propose SkillOS, an experience-driven RL training recipe for learning skill curation in self-evolving agents. SkillOS pairs a frozen agent executor that retrieves and applies skills with a trainable skill curator that updates an external SkillRepo from accumulated experience. To provide learning signals for curation, we design composite rewards and train on grouped task streams based on skill-relevant task dependencies, where earlier trajectories update the SkillRepo, and later related tasks evaluate these updates. Across multi-turn agentic tasks and single-turn reasoning tasks, SkillOS consistently outperforms memory-free and strong memory-based baselines in both effectiveness and efficiency, with the learned skill curator generalizing across different executor backbones and task domains. Further analyses show that the learned curator produces more targeted skill use, while the skills in SkillRepo evolve into more richly structured Markdown files that encode higher-level meta-skills over time.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06614
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SkillOS: Learning Skill Curation for Self-Evolving Agents
Ouyang, Siru
Yan, Jun
Chen, Yanfei
Han, Rujun
Wang, Zifeng
Mishra, Bhavana Dalvi
Meng, Rui
Li, Chun-Liang
Jiao, Yizhu
Zha, Kaiwen
Shen, Maohao
Tirumalashetty, Vishy
Lee, George
Han, Jiawei
Pfister, Tomas
Lee, Chen-Yu
Artificial Intelligence
Computation and Language
LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate for self-evolution, where high-quality skill curation serves as the key bottleneck. Existing approaches either rely on manual skill curation, prescribe heuristic skill operations, or train for short-horizon skill operations. However, they still struggle to learn complex long-term curation policies from indirect and delayed feedback. To tackle this challenge, we propose SkillOS, an experience-driven RL training recipe for learning skill curation in self-evolving agents. SkillOS pairs a frozen agent executor that retrieves and applies skills with a trainable skill curator that updates an external SkillRepo from accumulated experience. To provide learning signals for curation, we design composite rewards and train on grouped task streams based on skill-relevant task dependencies, where earlier trajectories update the SkillRepo, and later related tasks evaluate these updates. Across multi-turn agentic tasks and single-turn reasoning tasks, SkillOS consistently outperforms memory-free and strong memory-based baselines in both effectiveness and efficiency, with the learned skill curator generalizing across different executor backbones and task domains. Further analyses show that the learned curator produces more targeted skill use, while the skills in SkillRepo evolve into more richly structured Markdown files that encode higher-level meta-skills over time.
title SkillOS: Learning Skill Curation for Self-Evolving Agents
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.06614