Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Xiaoyu, Zhao, Sitong, Wang, Haotian, Chen, Shuaiting, Peng, Yiping, Ji, Yunjie, Zhao, Han, Li, Xiangang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918008686903296
author Tian, Xiaoyu
Zhao, Sitong
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Ji, Yunjie
Zhao, Han
Li, Xiangang
author_facet Tian, Xiaoyu
Zhao, Sitong
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Ji, Yunjie
Zhao, Han
Li, Xiangang
contents Despite significant advances in long-context reasoning by large language models (LLMs), primarily through Online Reinforcement Learning (RL) methods, these approaches incur substantial computational costs and complexity. In contrast, simpler and more economical Offline RL methods remain underexplored. To address this gap, we investigate the effectiveness of Offline RL methods, specifically Direct Preference Optimization (DPO) and its length-desensitized variant LD-DPO, in enhancing the reasoning capabilities of LLMs. Extensive experiments across multiple reasoning benchmarks demonstrate that these simpler Offline RL methods substantially improve model performance, achieving an average enhancement of 3.3\%, with a particularly notable increase of 10.1\% on the challenging Arena-Hard benchmark. Furthermore, we analyze DPO's sensitivity to output length, emphasizing that increasing reasoning length should align with semantic richness, as indiscriminate lengthening may adversely affect model performance. We provide comprehensive descriptions of our data processing and training methodologies, offering empirical evidence and practical insights for developing more cost-effective Offline RL approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
Tian, Xiaoyu
Zhao, Sitong
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Ji, Yunjie
Zhao, Han
Li, Xiangang
Computation and Language
Despite significant advances in long-context reasoning by large language models (LLMs), primarily through Online Reinforcement Learning (RL) methods, these approaches incur substantial computational costs and complexity. In contrast, simpler and more economical Offline RL methods remain underexplored. To address this gap, we investigate the effectiveness of Offline RL methods, specifically Direct Preference Optimization (DPO) and its length-desensitized variant LD-DPO, in enhancing the reasoning capabilities of LLMs. Extensive experiments across multiple reasoning benchmarks demonstrate that these simpler Offline RL methods substantially improve model performance, achieving an average enhancement of 3.3\%, with a particularly notable increase of 10.1\% on the challenging Arena-Hard benchmark. Furthermore, we analyze DPO's sensitivity to output length, emphasizing that increasing reasoning length should align with semantic richness, as indiscriminate lengthening may adversely affect model performance. We provide comprehensive descriptions of our data processing and training methodologies, offering empirical evidence and practical insights for developing more cost-effective Offline RL approaches.
title Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
topic Computation and Language
url https://arxiv.org/abs/2505.02142