Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Xiaoyu, Zhao, Sitong, Wang, Haotian, Chen, Shuaiting, Peng, Yiping, Ji, Yunjie, Zhao, Han, Li, Xiangang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
by: Ji, Yunjie, et al.
Published: (2025)
by: Ji, Yunjie, et al.
Published: (2025)
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
by: Wang, Haotian, et al.
Published: (2025)
by: Wang, Haotian, et al.
Published: (2025)
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
by: Ji, Yunjie, et al.
Published: (2025)
by: Ji, Yunjie, et al.
Published: (2025)
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
Not All Correct Answers Are Equal: Why Your Distillation Source Matters
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
by: Xu, Ruoxi, et al.
Published: (2025)
by: Xu, Ruoxi, et al.
Published: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny
by: Yan, Chuanhao, et al.
Published: (2025)
by: Yan, Chuanhao, et al.
Published: (2025)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
by: Chen, Mengyang, et al.
Published: (2023)
by: Chen, Mengyang, et al.
Published: (2023)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
by: Wen, Cheng, et al.
Published: (2025)
by: Wen, Cheng, et al.
Published: (2025)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
by: Ning, Xuefei, et al.
Published: (2024)
by: Ning, Xuefei, et al.
Published: (2024)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
by: Wang, Sijie, et al.
Published: (2025)
by: Wang, Sijie, et al.
Published: (2025)
DISA: Offline Importance Sampling for Distribution-Matching LLM-RL
by: Wang, Shaobo, et al.
Published: (2026)
by: Wang, Shaobo, et al.
Published: (2026)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
by: Tong, Zhao, et al.
Published: (2026)
by: Tong, Zhao, et al.
Published: (2026)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
by: Chen, Zui, et al.
Published: (2024)
by: Chen, Zui, et al.
Published: (2024)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets
by: He, Mingqian, et al.
Published: (2025)
by: He, Mingqian, et al.
Published: (2025)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Bridging Offline and Online Reinforcement Learning for LLMs
by: Lanchantin, Jack, et al.
Published: (2025)
by: Lanchantin, Jack, et al.
Published: (2025)
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
by: Wang, Tianle, et al.
Published: (2026)
by: Wang, Tianle, et al.
Published: (2026)
LexPro-1.0 Technical Report
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
by: Pournemat, Mobina, et al.
Published: (2025)
by: Pournemat, Mobina, et al.
Published: (2025)
Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance
by: Qian, Lingfei, et al.
Published: (2025)
by: Qian, Lingfei, et al.
Published: (2025)
ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas
by: Tian, Xiaoyu, et al.
Published: (2026)
by: Tian, Xiaoyu, et al.
Published: (2026)
MoL-RL: Distilling Multi-Step Environmental Feedback into LLMs for Feedback-Independent Reasoning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
by: Li, Bingru, et al.
Published: (2024)
by: Li, Bingru, et al.
Published: (2024)
Reasoning with Graphs: Structuring Implicit Knowledge to Enhance LLMs Reasoning
by: Han, Haoyu, et al.
Published: (2025)
by: Han, Haoyu, et al.
Published: (2025)
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
by: Qiu, Pengcheng, et al.
Published: (2025)
by: Qiu, Pengcheng, et al.
Published: (2025)
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
by: Yuan, Danlong, et al.
Published: (2025)
by: Yuan, Danlong, et al.
Published: (2025)
ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-Thought
by: Wang, Fanmeng, et al.
Published: (2026)
by: Wang, Fanmeng, et al.
Published: (2026)
Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
by: Chen, Xinghao, et al.
Published: (2025)
by: Chen, Xinghao, et al.
Published: (2025)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
by: Lei, Fangyu, et al.
Published: (2025)
by: Lei, Fangyu, et al.
Published: (2025)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
by: Kim, Jisu, et al.
Published: (2025)
by: Kim, Jisu, et al.
Published: (2025)
Similar Items
-
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
by: Ji, Yunjie, et al.
Published: (2025) -
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
by: Wang, Haotian, et al.
Published: (2025) -
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
by: Ji, Yunjie, et al.
Published: (2025) -
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
by: Tian, Xiaoyu, et al.
Published: (2025) -
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training
by: Zhao, Han, et al.
Published: (2025)