Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shi, Taiwei, Wu, Yiyang, Song, Linxin, Zhou, Tianyi, Zhao, Jieyu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Hallucination Tax of Reinforcement Finetuning
par: Song, Linxin, et autres
Publié: (2025)
par: Song, Linxin, et autres
Publié: (2025)
Experiential Reinforcement Learning
par: Shi, Taiwei, et autres
Publié: (2026)
par: Shi, Taiwei, et autres
Publié: (2026)
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
par: Li, Zongqian, et autres
Publié: (2026)
par: Li, Zongqian, et autres
Publié: (2026)
Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base
par: Song, Linxin, et autres
Publié: (2025)
par: Song, Linxin, et autres
Publié: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
par: Shi, Taiwei, et autres
Publié: (2023)
par: Shi, Taiwei, et autres
Publié: (2023)
Video-Based Reward Modeling for Computer-Use Agents
par: Song, Linxin, et autres
Publié: (2026)
par: Song, Linxin, et autres
Publié: (2026)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
par: Indurthi, Sathish Reddy, et autres
Publié: (2024)
par: Indurthi, Sathish Reddy, et autres
Publié: (2024)
Vanishing Gradients in Reinforcement Finetuning of Language Models
par: Razin, Noam, et autres
Publié: (2023)
par: Razin, Noam, et autres
Publié: (2023)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
par: Shen, Yiyang, et autres
Publié: (2026)
par: Shen, Yiyang, et autres
Publié: (2026)
OMoE: Diversifying Mixture of Low-Rank Adaptation by Orthogonal Finetuning
par: Feng, Jinyuan, et autres
Publié: (2025)
par: Feng, Jinyuan, et autres
Publié: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
par: Liu, Hanbing, et autres
Publié: (2025)
par: Liu, Hanbing, et autres
Publié: (2025)
Explaining Length Bias in LLM-Based Preference Evaluations
par: Hu, Zhengyu, et autres
Publié: (2024)
par: Hu, Zhengyu, et autres
Publié: (2024)
Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation
par: Pan, Yijun, et autres
Publié: (2025)
par: Pan, Yijun, et autres
Publié: (2025)
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
par: Liang, Jing, et autres
Publié: (2025)
par: Liang, Jing, et autres
Publié: (2025)
Transfer Learning for Finetuning Large Language Models
par: Strangmann, Tobias, et autres
Publié: (2024)
par: Strangmann, Tobias, et autres
Publié: (2024)
Prompt Curriculum Learning for Efficient LLM Post-Training
par: Gao, Zhaolin, et autres
Publié: (2025)
par: Gao, Zhaolin, et autres
Publié: (2025)
AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting
par: Li, Renda, et autres
Publié: (2025)
par: Li, Renda, et autres
Publié: (2025)
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
par: Jiang, Guochao, et autres
Publié: (2025)
par: Jiang, Guochao, et autres
Publié: (2025)
Why Softmax Attention Outperforms Linear Attention
par: Deng, Yichuan, et autres
Publié: (2023)
par: Deng, Yichuan, et autres
Publié: (2023)
MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
par: Page, Saurabh, et autres
Publié: (2025)
par: Page, Saurabh, et autres
Publié: (2025)
KnowLA: Enhancing Parameter-efficient Finetuning with Knowledgeable Adaptation
par: Luo, Xindi, et autres
Publié: (2024)
par: Luo, Xindi, et autres
Publié: (2024)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
par: Longpre, Shayne, et autres
Publié: (2025)
par: Longpre, Shayne, et autres
Publié: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
par: Zhang, Biao, et autres
Publié: (2024)
par: Zhang, Biao, et autres
Publié: (2024)
From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
par: Zhao, Eric, et autres
Publié: (2025)
par: Zhao, Eric, et autres
Publié: (2025)
Learning Dynamics of LLM Finetuning
par: Ren, Yi, et autres
Publié: (2024)
par: Ren, Yi, et autres
Publié: (2024)
PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
par: Hayou, Soufiane, et autres
Publié: (2025)
par: Hayou, Soufiane, et autres
Publié: (2025)
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
par: Feng, Jiahai, et autres
Publié: (2024)
par: Feng, Jiahai, et autres
Publié: (2024)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
par: Gupta, Prakhar, et autres
Publié: (2026)
par: Gupta, Prakhar, et autres
Publié: (2026)
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
par: Liu, Weiyang, et autres
Publié: (2023)
par: Liu, Weiyang, et autres
Publié: (2023)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
par: Wang, Xiyao, et autres
Publié: (2024)
par: Wang, Xiyao, et autres
Publié: (2024)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
par: Xu, Zhuoyan, et autres
Publié: (2024)
par: Xu, Zhuoyan, et autres
Publié: (2024)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
par: Guo, Yiju, et autres
Publié: (2026)
par: Guo, Yiju, et autres
Publié: (2026)
Internalizing World Models via Self-Play Finetuning for Agentic RL
par: Chen, Shiqi, et autres
Publié: (2025)
par: Chen, Shiqi, et autres
Publié: (2025)
Improving Sparse Memory Finetuning
par: Goyal, Satyam, et autres
Publié: (2026)
par: Goyal, Satyam, et autres
Publié: (2026)
Predicting Emergent Capabilities by Finetuning
par: Snell, Charlie, et autres
Publié: (2024)
par: Snell, Charlie, et autres
Publié: (2024)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
par: Qin, Haotong, et autres
Publié: (2024)
par: Qin, Haotong, et autres
Publié: (2024)
ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
par: Feng, Zihao, et autres
Publié: (2025)
par: Feng, Zihao, et autres
Publié: (2025)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
par: Huang, Yangyi, et autres
Publié: (2026)
par: Huang, Yangyi, et autres
Publié: (2026)
Time Sensitive Knowledge Editing through Efficient Finetuning
par: Ge, Xiou, et autres
Publié: (2024)
par: Ge, Xiou, et autres
Publié: (2024)
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
par: Zeng, Zhiyuan, et autres
Publié: (2025)
par: Zeng, Zhiyuan, et autres
Publié: (2025)
Documents similaires
-
The Hallucination Tax of Reinforcement Finetuning
par: Song, Linxin, et autres
Publié: (2025) -
Experiential Reinforcement Learning
par: Shi, Taiwei, et autres
Publié: (2026) -
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
par: Li, Zongqian, et autres
Publié: (2026) -
Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base
par: Song, Linxin, et autres
Publié: (2025) -
Safer-Instruct: Aligning Language Models with Automated Preference Data
par: Shi, Taiwei, et autres
Publié: (2023)