Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Zhenwen, Zhou, Yujun, Lu, Sidi, Zhang, Xiangliang, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)
von: Yu, Dian, et al.
Veröffentlicht: (2025)
Causally-Enhanced Reinforcement Policy Optimization
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
von: Wang, Xiangqi, et al.
Veröffentlicht: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
Capability-Oriented Training Induced Alignment Risk
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
Manipulating Predictions over Discrete Inputs in Machine Teaching
von: Wu, Xiaodong, et al.
Veröffentlicht: (2024)
von: Wu, Xiaodong, et al.
Veröffentlicht: (2024)
Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
von: Wang, Wenkai, et al.
Veröffentlicht: (2026)
von: Wang, Wenkai, et al.
Veröffentlicht: (2026)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
Zero-Shot Relational Learning for Multimodal Knowledge Graphs
von: Cai, Rui, et al.
Veröffentlicht: (2024)
von: Cai, Rui, et al.
Veröffentlicht: (2024)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024)
von: Ge, Tao, et al.
Veröffentlicht: (2024)
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
von: Tang, Mohan, et al.
Veröffentlicht: (2026)
von: Tang, Mohan, et al.
Veröffentlicht: (2026)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Mingxuan, et al.
Veröffentlicht: (2025)
TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction
von: Xu, Tommy, et al.
Veröffentlicht: (2025)
von: Xu, Tommy, et al.
Veröffentlicht: (2025)
Learning to Clean: Reinforcement Learning for Noisy Label Correction
von: Heidari, Marzi, et al.
Veröffentlicht: (2025)
von: Heidari, Marzi, et al.
Veröffentlicht: (2025)
Counterfactual Explanations for Continuous Action Reinforcement Learning
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
von: Dong, Shuyang, et al.
Veröffentlicht: (2025)
Edge Contrastive Learning: An Augmentation-Free Graph Contrastive Learning Model
von: Li, Yujun, et al.
Veröffentlicht: (2024)
von: Li, Yujun, et al.
Veröffentlicht: (2024)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
FairGRPO: Fair Reinforcement Learning for Equitable Clinical Reasoning
von: Dai, Shiqi, et al.
Veröffentlicht: (2025)
von: Dai, Shiqi, et al.
Veröffentlicht: (2025)
Learning to Reason Efficiently with Discounted Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs
von: Zeng, Xinyue, et al.
Veröffentlicht: (2026)
von: Zeng, Xinyue, et al.
Veröffentlicht: (2026)
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
von: Shi, Yucheng, et al.
Veröffentlicht: (2026)
von: Shi, Yucheng, et al.
Veröffentlicht: (2026)
SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
von: Liu, Huanyu, et al.
Veröffentlicht: (2025)
von: Liu, Huanyu, et al.
Veröffentlicht: (2025)
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
von: Lang, Yicheng, et al.
Veröffentlicht: (2025)
von: Lang, Yicheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025) -
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026) -
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026) -
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025) -
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)