How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Simon, Kaur, Simran, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024)
by: Kaur, Simran, et al.
Published: (2024)
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
Prioritized Replay for RL Post-training
by: Fatemi, Mehdi
Published: (2026)
by: Fatemi, Mehdi
Published: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
On the Power of Context-Enhanced Learning in LLMs
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Optimistic Verifiable Training by Controlling Hardware Nondeterminism
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025)
by: Kempf, Elias, et al.
Published: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
CRePE: Convolution-aware Relative Importance in Post-training Pruning with Efficient Search
by: Park, Cheonjun
Published: (2026)
by: Park, Cheonjun
Published: (2026)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
The unregulated plant‐based ‘milk’ industry: A threat to nutrition, health and safety?
by: Simran Kaur Arora
Published: (2024)
by: Simran Kaur Arora
Published: (2024)
On the Impossibility of Retrain Equivalence in Machine Unlearning
by: Yu, Jiatong, et al.
Published: (2025)
by: Yu, Jiatong, et al.
Published: (2025)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
by: Xia, Peng, et al.
Published: (2026)
by: Xia, Peng, et al.
Published: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Skill Reuse as Compression in Agentic RL
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
by: Zabounidis, Renos, et al.
Published: (2026)
by: Zabounidis, Renos, et al.
Published: (2026)
A Quadratic Synchronization Rule for Distributed Deep Learning
by: Gu, Xinran, et al.
Published: (2023)
by: Gu, Xinran, et al.
Published: (2023)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
by: He, Zelin, et al.
Published: (2026)
by: He, Zelin, et al.
Published: (2026)
ASAP: Unsupervised Post-training with Label Distribution Shift Adaptive Learning Rate
by: Park, Heewon, et al.
Published: (2025)
by: Park, Heewon, et al.
Published: (2025)
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs
by: Jha, Piyush, et al.
Published: (2024)
by: Jha, Piyush, et al.
Published: (2024)
SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning
by: Zhang, Ruiqi, et al.
Published: (2025)
by: Zhang, Ruiqi, et al.
Published: (2025)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
by: Wu, Haoze, et al.
Published: (2025)
by: Wu, Haoze, et al.
Published: (2025)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
by: Lu, Aojun, et al.
Published: (2026)
by: Lu, Aojun, et al.
Published: (2026)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
by: Wu, Yikai, et al.
Published: (2025)
by: Wu, Yikai, et al.
Published: (2025)
Learning Dynamics in RL Post-Training for Language Models
by: Tomihari, Akiyoshi
Published: (2026)
by: Tomihari, Akiyoshi
Published: (2026)
RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
by: Samineni, Soumya Rani, et al.
Published: (2025)
by: Samineni, Soumya Rani, et al.
Published: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Similar Items
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
by: Kaur, Simran, et al.
Published: (2024) -
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024) -
Skill-Targeted Adaptive Training
by: He, Yinghui, et al.
Published: (2025) -
Prioritized Replay for RL Post-training
by: Fatemi, Mehdi
Published: (2026) -
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
by: Khalifa, Muhammad, et al.
Published: (2026)