Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Lu, Liang, Hao, Qiang, Meiyi, Tang, Lexiang, Ma, Xiaochen, Wong, Zhen Hao, Niu, Junbo, Shen, Chengyu, He, Runming, Li, Yanhao, Cui, Bin, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let's Verify Math Questions Step by Step
by: Shen, Chengyu, et al.
Published: (2025)
by: Shen, Chengyu, et al.
Published: (2025)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026)
by: Shen, Chengyu, et al.
Published: (2026)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
by: Li, Yanhao, et al.
Published: (2025)
by: Li, Yanhao, et al.
Published: (2025)
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization
by: Zhao, Zhengyang, et al.
Published: (2026)
by: Zhao, Zhengyang, et al.
Published: (2026)
Online Learning from Strategic Human Feedback in LLM Fine-Tuning
by: Hao, Shugang, et al.
Published: (2024)
by: Hao, Shugang, et al.
Published: (2024)
IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Supervised Fine-Tuning as Inverse Reinforcement Learning
by: Sun, Hao
Published: (2024)
by: Sun, Hao
Published: (2024)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Formal Logic Enabled Personalized Federated Learning Through Property Inference
by: An, Ziyan, et al.
Published: (2024)
by: An, Ziyan, et al.
Published: (2024)
Autocrats Can't Always Get What They Want
by: Brown, Nathan J., et al.
Published: (2024)
by: Brown, Nathan J., et al.
Published: (2024)
DARO: Difficulty-Aware Reweighting Policy Optimization
by: Zhou, Jingyu, et al.
Published: (2025)
by: Zhou, Jingyu, et al.
Published: (2025)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
They Can't Always Find What They Want; Kids' Online Behaviors Have Researchers Scratching Their Heads
by: Minkel, Walter
Published: (2004)
by: Minkel, Walter
Published: (2004)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Gradual Learning: Optimizing Fine-Tuning with Partially Mastered Knowledge in Large Language Models
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
Can’t Touch This
Published: (2024)
Published: (2024)
Why I Can't Create a Learning Center
by: Miller, Rosalind
Published: (1975)
by: Miller, Rosalind
Published: (1975)
Learning with Preserving for Continual Multitask Learning
by: Wang, Hanchen David, et al.
Published: (2025)
by: Wang, Hanchen David, et al.
Published: (2025)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
Hardest Monotone Functions for Evolutionary Algorithms
by: Kaufmann, Marc, et al.
Published: (2023)
by: Kaufmann, Marc, et al.
Published: (2023)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Human Resilience in the AI Era -- What Machines Can't Replace
by: Liu, Shaoshan, et al.
Published: (2025)
by: Liu, Shaoshan, et al.
Published: (2025)
AbsenceBench: Language Models Can't Tell What's Missing
by: Fu, Harvey Yiyun, et al.
Published: (2025)
by: Fu, Harvey Yiyun, et al.
Published: (2025)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Integrating Offline Pre-Training with Online Fine-Tuning: A Reinforcement Learning Approach for Robot Social Navigation
by: Su, Run, et al.
Published: (2025)
by: Su, Run, et al.
Published: (2025)
They Can't Hear Us Does Not Mean We Can't Serve Them.
by: McDaniel, Julie Ann
Published: (1992)
by: McDaniel, Julie Ann
Published: (1992)
QAEncoder: Towards Aligned Representation Learning in Question Answering Systems
by: Wang, Zhengren, et al.
Published: (2024)
by: Wang, Zhengren, et al.
Published: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
by: Upadhyay, Ujjwal, et al.
Published: (2025)
by: Upadhyay, Ujjwal, et al.
Published: (2025)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
by: Li, Xiangyang, et al.
Published: (2025)
by: Li, Xiangyang, et al.
Published: (2025)
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
by: Ma, Hao, et al.
Published: (2024)
by: Ma, Hao, et al.
Published: (2024)
Equal Requests are Asymptotically Hardest for Data Recovery
by: Lember, Jüri, et al.
Published: (2024)
by: Lember, Jüri, et al.
Published: (2024)
Similar Items
-
Let's Verify Math Questions Step by Step
by: Shen, Chengyu, et al.
Published: (2025) -
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
by: Liang, Hao, et al.
Published: (2026) -
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026) -
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
by: Wong, Zhen Hao, et al.
Published: (2025) -
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
by: Wong, Zhen Hao, et al.
Published: (2025)