Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yifan, Shen, Jingyan, Wang, Yibin, Chen, Tianyu, Wang, Zhendong, Zhou, Mingyuan, Zhang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
2D-OOB: Attributing Data Contribution Through Joint Valuation Framework
by: Sun, Yifan, et al.
Published: (2024)
by: Sun, Yifan, et al.
Published: (2024)
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Influence-Preserving Proxies for Gradient-Based Data Selection in LLM Fine-tuning
by: Chen, Sirui, et al.
Published: (2026)
by: Chen, Sirui, et al.
Published: (2026)
Enhancing and Accelerating Diffusion-Based Inverse Problem Solving through Measurements Optimization
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
by: Liu, Xiangyan, et al.
Published: (2025)
by: Liu, Xiangyan, et al.
Published: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
by: Wang, Qingzheng, et al.
Published: (2025)
by: Wang, Qingzheng, et al.
Published: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Score Distillation Beyond Acceleration: Generative Modeling from Corrupted Data
by: Zhang, Yasi, et al.
Published: (2025)
by: Zhang, Yasi, et al.
Published: (2025)
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
by: Shen, Han, et al.
Published: (2024)
by: Shen, Han, et al.
Published: (2024)
Token-level Data Selection for Safe LLM Fine-tuning
by: Li, Yanping, et al.
Published: (2026)
by: Li, Yanping, et al.
Published: (2026)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
by: Zhou, Mingyuan, et al.
Published: (2024)
by: Zhou, Mingyuan, et al.
Published: (2024)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
by: Wang, Fangxin, et al.
Published: (2026)
by: Wang, Fangxin, et al.
Published: (2026)
Data-efficient Fine-tuning for LLM-based Recommendation
by: Lin, Xinyu, et al.
Published: (2024)
by: Lin, Xinyu, et al.
Published: (2024)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Reinforcement Fine-Tuning for Materials Design
by: Cao, Zhendong, et al.
Published: (2025)
by: Cao, Zhendong, et al.
Published: (2025)
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
by: Song, Yueqi, et al.
Published: (2025)
by: Song, Yueqi, et al.
Published: (2025)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
by: Wang, Qibin, et al.
Published: (2024)
by: Wang, Qibin, et al.
Published: (2024)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Data-efficient LLM Fine-tuning for Code Generation
by: Lv, Weijie, et al.
Published: (2025)
by: Lv, Weijie, et al.
Published: (2025)
Pricing Online LLM Services with Data-Calibrated Stackelberg Routing Game
by: Guo, Zhendong, et al.
Published: (2025)
by: Guo, Zhendong, et al.
Published: (2025)
Few-Step Diffusion via Score identity Distillation
by: Zhou, Mingyuan, et al.
Published: (2025)
by: Zhou, Mingyuan, et al.
Published: (2025)
Fine-tuning for Data-enabled Predictive Control of Noisy Systems by Reinforcement Learning
by: Wang, Jinbao, et al.
Published: (2025)
by: Wang, Jinbao, et al.
Published: (2025)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud
by: Yue, Yuanhao, et al.
Published: (2024)
by: Yue, Yuanhao, et al.
Published: (2024)
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
by: Zhang, Tonghe, et al.
Published: (2025)
by: Zhang, Tonghe, et al.
Published: (2025)
Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization
by: Chen, Xuxi, et al.
Published: (2024)
by: Chen, Xuxi, et al.
Published: (2024)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
Prioritized Trajectory Replay: A Replay Memory for Data-driven Reinforcement Learning
by: Liu, Jinyi, et al.
Published: (2023)
by: Liu, Jinyi, et al.
Published: (2023)
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
by: Wu, Jianghao, et al.
Published: (2026)
by: Wu, Jianghao, et al.
Published: (2026)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
A Novel Framework for Online Supervised Learning with Feature Selection
by: Sun, Lizhe, et al.
Published: (2018)
by: Sun, Lizhe, et al.
Published: (2018)
Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
by: Saroufim, Mark, et al.
Published: (2025)
by: Saroufim, Mark, et al.
Published: (2025)
Similar Items
-
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning
by: Chen, Tianyu, et al.
Published: (2024) -
2D-OOB: Attributing Data Contribution Through Joint Valuation Framework
by: Sun, Yifan, et al.
Published: (2024) -
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
by: Surana, Rohan, et al.
Published: (2026) -
Influence-Preserving Proxies for Gradient-Based Data Selection in LLM Fine-tuning
by: Chen, Sirui, et al.
Published: (2026) -
Enhancing and Accelerating Diffusion-Based Inverse Problem Solving through Measurements Optimization
by: Chen, Tianyu, et al.
Published: (2024)