Dynamic Dual-Granularity Skill Bank for Agentic RL
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911715084468224 |
|---|---|
| author | Tu, Songjun Xu, Chengdong Zhang, Qichao Zhang, Yaocheng Lan, Xiangyuan Li, Linjing Li, Dong Zhao, Dongbin |
| author_facet | Tu, Songjun Xu, Chengdong Zhang, Qichao Zhang, Yaocheng Lan, Xiangyuan Li, Linjing Li, Dong Zhao, Dongbin |
| contents | Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_28716 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Dynamic Dual-Granularity Skill Bank for Agentic RL Tu, Songjun Xu, Chengdong Zhang, Qichao Zhang, Yaocheng Lan, Xiangyuan Li, Linjing Li, Dong Zhao, Dongbin Artificial Intelligence 68T05 I.2.6; I.2.11 Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead. |
| title | Dynamic Dual-Granularity Skill Bank for Agentic RL |
| topic | Artificial Intelligence 68T05 I.2.6; I.2.11 |
| url | https://arxiv.org/abs/2603.28716 |