COPO: Consistency-Aware Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Jinghang, Chen, Jiawei, Shao, Hang, Ma, Hao, Li, Mingcheng, Shen, Xintian, Zheng, Lihao, Chen, Wei, Wei, Tao, Zhang, Lihua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911093883928576
author Han, Jinghang
Chen, Jiawei
Shao, Hang
Ma, Hao
Li, Mingcheng
Shen, Xintian
Zheng, Lihao
Chen, Wei
Wei, Tao
Zhang, Lihua
author_facet Han, Jinghang
Chen, Jiawei
Shao, Hang
Ma, Hao
Li, Mingcheng
Shen, Xintian
Zheng, Lihao
Chen, Wei
Wei, Tao
Zhang, Lihua
contents Reinforcement learning has significantly enhanced the reasoning capabilities of Large Language Models (LLMs) in complex problem-solving tasks. Recently, the introduction of DeepSeek R1 has inspired a surge of interest in leveraging rule-based rewards as a low-cost alternative for computing advantage functions and guiding policy optimization. However, a common challenge observed across many replication and extension efforts is that when multiple sampled responses under a single prompt converge to identical outcomes, whether correct or incorrect, the group-based advantage degenerates to zero. This leads to vanishing gradients and renders the corresponding samples ineffective for learning, ultimately limiting training efficiency and downstream performance. To address this issue, we propose a consistency-aware policy optimization framework that introduces a structured global reward based on outcome consistency, the global loss based on it ensures that, even when model outputs show high intra-group consistency, the training process still receives meaningful learning signals, which encourages the generation of correct and self-consistent reasoning paths from a global perspective. Furthermore, we incorporate an entropy-based soft blending mechanism that adaptively balances local advantage estimation with global optimization, enabling dynamic transitions between exploration and convergence throughout training. Our method introduces several key innovations in both reward design and optimization strategy. We validate its effectiveness through substantial performance gains on multiple mathematical reasoning benchmarks, highlighting the proposed framework's robustness and general applicability. Code of this work has been released at https://github.com/hijih/copo-code.git.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COPO: Consistency-Aware Policy Optimization
Han, Jinghang
Chen, Jiawei
Shao, Hang
Ma, Hao
Li, Mingcheng
Shen, Xintian
Zheng, Lihao
Chen, Wei
Wei, Tao
Zhang, Lihua
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement learning has significantly enhanced the reasoning capabilities of Large Language Models (LLMs) in complex problem-solving tasks. Recently, the introduction of DeepSeek R1 has inspired a surge of interest in leveraging rule-based rewards as a low-cost alternative for computing advantage functions and guiding policy optimization. However, a common challenge observed across many replication and extension efforts is that when multiple sampled responses under a single prompt converge to identical outcomes, whether correct or incorrect, the group-based advantage degenerates to zero. This leads to vanishing gradients and renders the corresponding samples ineffective for learning, ultimately limiting training efficiency and downstream performance. To address this issue, we propose a consistency-aware policy optimization framework that introduces a structured global reward based on outcome consistency, the global loss based on it ensures that, even when model outputs show high intra-group consistency, the training process still receives meaningful learning signals, which encourages the generation of correct and self-consistent reasoning paths from a global perspective. Furthermore, we incorporate an entropy-based soft blending mechanism that adaptively balances local advantage estimation with global optimization, enabling dynamic transitions between exploration and convergence throughout training. Our method introduces several key innovations in both reward design and optimization strategy. We validate its effectiveness through substantial performance gains on multiple mathematical reasoning benchmarks, highlighting the proposed framework's robustness and general applicability. Code of this work has been released at https://github.com/hijih/copo-code.git.
title COPO: Consistency-Aware Policy Optimization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.04138