Skill Reuse as Compression in Agentic RL
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhikun, Feng, Yu, Dineen, Jacob, Shi, Taiwei, Zhao, Jieyu, Zhou, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Experiential Reinforcement Learning
by: Shi, Taiwei, et al.
Published: (2026)
by: Shi, Taiwei, et al.
Published: (2026)
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
by: He, Zelin, et al.
Published: (2026)
by: He, Zelin, et al.
Published: (2026)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026)
by: Dineen, Jacob, et al.
Published: (2026)
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
by: Gao, Zijun, et al.
Published: (2025)
by: Gao, Zijun, et al.
Published: (2025)
Agentic Skill Discovery
by: Zhao, Xufeng, et al.
Published: (2024)
by: Zhao, Xufeng, et al.
Published: (2024)
Vision-Language Model Selection and Reuse for Downstream Adaptation
by: Tan, Hao-Zhe, et al.
Published: (2025)
by: Tan, Hao-Zhe, et al.
Published: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
by: Shi, Taiwei, et al.
Published: (2023)
by: Shi, Taiwei, et al.
Published: (2023)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
by: Dai, Weinan, et al.
Published: (2026)
by: Dai, Weinan, et al.
Published: (2026)
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
by: Liu, Tong, et al.
Published: (2026)
by: Liu, Tong, et al.
Published: (2026)
AB-Cache: Training-Free Acceleration of Diffusion Models via Adams-Bashforth Cached Feature Reuse
by: Yu, Zichao, et al.
Published: (2025)
by: Yu, Zichao, et al.
Published: (2025)
Offline RLAIF: Piloting VLM Feedback for RL via SFO
by: Beck, Jacob
Published: (2025)
by: Beck, Jacob
Published: (2025)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
by: Feng, Xiao, et al.
Published: (2026)
by: Feng, Xiao, et al.
Published: (2026)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
by: Wang, Kunyun, et al.
Published: (2025)
by: Wang, Kunyun, et al.
Published: (2025)
Tiled Bit Networks: Sub-Bit Neural Network Compression Through Reuse of Learnable Binary Vectors
by: Gorbett, Matt, et al.
Published: (2024)
by: Gorbett, Matt, et al.
Published: (2024)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
by: Guan, Zhong, et al.
Published: (2026)
by: Guan, Zhong, et al.
Published: (2026)
Agentic Proposing: Enhancing Large Language Model Reasoning via Compositional Skill Synthesis
by: Jiao, Zhengbo, et al.
Published: (2026)
by: Jiao, Zhengbo, et al.
Published: (2026)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
by: Setlur, Amrith, et al.
Published: (2026)
by: Setlur, Amrith, et al.
Published: (2026)
CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
by: Mai, Shinji, et al.
Published: (2025)
by: Mai, Shinji, et al.
Published: (2025)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
by: Tang, Pingzhi, et al.
Published: (2026)
by: Tang, Pingzhi, et al.
Published: (2026)
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
by: Miao, Yuchun, et al.
Published: (2026)
by: Miao, Yuchun, et al.
Published: (2026)
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
by: Yu, Haonan, et al.
Published: (2025)
by: Yu, Haonan, et al.
Published: (2025)
Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression
by: Tian, Ye, et al.
Published: (2026)
by: Tian, Ye, et al.
Published: (2026)
TWISTED-RL: Hierarchical Skilled Agents for Knot-Tying without Human Demonstrations
by: Freund, Guy, et al.
Published: (2026)
by: Freund, Guy, et al.
Published: (2026)
VendiRL: A Framework for Self-Supervised Reinforcement Learning of Diversely Diverse Skills
by: Lintunen, Erik M.
Published: (2025)
by: Lintunen, Erik M.
Published: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
by: Pan, Yijun, et al.
Published: (2025)
by: Pan, Yijun, et al.
Published: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Analytical Lyapunov Function Discovery: An RL-based Generative Approach
by: Zou, Haohan, et al.
Published: (2025)
by: Zou, Haohan, et al.
Published: (2025)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
by: Shi, Jiang-Xin, et al.
Published: (2025)
by: Shi, Jiang-Xin, et al.
Published: (2025)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
DP-CSGP: Differentially Private Stochastic Gradient Push with Compressed Communication
by: Zhu, Zehan, et al.
Published: (2025)
by: Zhu, Zehan, et al.
Published: (2025)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
by: Zhao, Zibo, et al.
Published: (2026)
by: Zhao, Zibo, et al.
Published: (2026)
Augmenting Offline RL with Unlabeled Data
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Similar Items
-
Experiential Reinforcement Learning
by: Shi, Taiwei, et al.
Published: (2026) -
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
by: He, Zelin, et al.
Published: (2026) -
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026) -
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
by: Gao, Zijun, et al.
Published: (2025) -
Agentic Skill Discovery
by: Zhao, Xufeng, et al.
Published: (2024)