Gespeichert in:
| Hauptverfasser: | Lu, Sidi, Liang, Zhenwen, Ma, Dongyang, Wang, Yan, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.05085 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
Verified Critical Step Optimization for LLM Agents
von: Li, Mukai, et al.
Veröffentlicht: (2026)
von: Li, Mukai, et al.
Veröffentlicht: (2026)
Free(): Learning to Forget in Malloc-Only Reasoning Models
von: Zheng, Yilun, et al.
Veröffentlicht: (2026)
von: Zheng, Yilun, et al.
Veröffentlicht: (2026)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
von: Li, Mukai, et al.
Veröffentlicht: (2025)
von: Li, Mukai, et al.
Veröffentlicht: (2025)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
von: Ma, Junyu, et al.
Veröffentlicht: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
The End of Manual Decoding: Towards Truly End-to-End Language Models
von: Wang, Zhichao, et al.
Veröffentlicht: (2025)
von: Wang, Zhichao, et al.
Veröffentlicht: (2025)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data
von: Liang, Zhenwen, et al.
Veröffentlicht: (2026)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2026)
HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows
von: Yao, Wenlin, et al.
Veröffentlicht: (2024)
von: Yao, Wenlin, et al.
Veröffentlicht: (2024)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
von: He, Zhiwei, et al.
Veröffentlicht: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2024)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
von: Das, Souvik, et al.
Veröffentlicht: (2024)
von: Das, Souvik, et al.
Veröffentlicht: (2024)
DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
von: Liang, Tian, et al.
Veröffentlicht: (2025)
von: Liang, Tian, et al.
Veröffentlicht: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Inconsistent dialogue responses and how to recover from them
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024)
von: Ge, Tao, et al.
Veröffentlicht: (2024)
Using LLM to select the right SQL Query from candidates
von: Li, Zhenwen, et al.
Veröffentlicht: (2024)
von: Li, Zhenwen, et al.
Veröffentlicht: (2024)
Teaching LLMs to Refine with Tools
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhisong, et al.
Veröffentlicht: (2025)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
Don't Throw Away Your Pretrained Model
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)
von: Yu, Dian, et al.
Veröffentlicht: (2025)
Towards Solving More Challenging IMO Problems via Decoupled Reasoning and Proving
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026) -
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026) -
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025) -
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025) -
Verified Critical Step Optimization for LLM Agents
von: Li, Mukai, et al.
Veröffentlicht: (2026)