Causally-Enhanced Reinforcement Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiangqi, Huang, Yue, Zhou, Yujun, Luo, Xiaonan, Guo, Kehan, Zhang, Xiangliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911181281689600
author Wang, Xiangqi
Huang, Yue
Zhou, Yujun
Luo, Xiaonan
Guo, Kehan
Zhang, Xiangliang
author_facet Wang, Xiangqi
Huang, Yue
Zhou, Yujun
Luo, Xiaonan
Guo, Kehan
Zhang, Xiangliang
contents Large language models (LLMs) trained with reinforcement objectives often achieve superficially correct answers via shortcut strategies, pairing correct outputs with spurious or unfaithful reasoning and degrading under small causal perturbations. We introduce Causally-Enhanced Policy Optimization (CE-PO), a drop-in reward-shaping framework that augments policy optimization with a differentiable proxy for causal coherence along the generation pathway from prompt (Z) to rationale (X) to answer (Y). CE-PO estimates model-internal influence with Jacobian-based sensitivities, counterfactually hardens these signals to suppress nuisance cues, and fuses the resulting coherence score with task-accuracy feedback via a Minkowski (power-mean) combiner, exposing a single tunable between accuracy and coherence trade-off. The unified reward integrates with PPO/GRPO without architectural changes. Across reasoning benchmarks and causal stress tests, CE-PO reduces reward hacking and unfaithful chain-of-thought while improving robustness to correlation-causation flips and light counterfactual edits, all at near-parity accuracy. Experimental results across 4 datasets show that CE-PO improves accuracy over baselines by 5.49% on average (up to 9.58%), while improving robustness to correlation-causation flips and light counterfactual edits.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Causally-Enhanced Reinforcement Policy Optimization
Wang, Xiangqi
Huang, Yue
Zhou, Yujun
Luo, Xiaonan
Guo, Kehan
Zhang, Xiangliang
Machine Learning
Artificial Intelligence
Computation and Language
Large language models (LLMs) trained with reinforcement objectives often achieve superficially correct answers via shortcut strategies, pairing correct outputs with spurious or unfaithful reasoning and degrading under small causal perturbations. We introduce Causally-Enhanced Policy Optimization (CE-PO), a drop-in reward-shaping framework that augments policy optimization with a differentiable proxy for causal coherence along the generation pathway from prompt (Z) to rationale (X) to answer (Y). CE-PO estimates model-internal influence with Jacobian-based sensitivities, counterfactually hardens these signals to suppress nuisance cues, and fuses the resulting coherence score with task-accuracy feedback via a Minkowski (power-mean) combiner, exposing a single tunable between accuracy and coherence trade-off. The unified reward integrates with PPO/GRPO without architectural changes. Across reasoning benchmarks and causal stress tests, CE-PO reduces reward hacking and unfaithful chain-of-thought while improving robustness to correlation-causation flips and light counterfactual edits, all at near-parity accuracy. Experimental results across 4 datasets show that CE-PO improves accuracy over baselines by 5.49% on average (up to 9.58%), while improving robustness to correlation-causation flips and light counterfactual edits.
title Causally-Enhanced Reinforcement Policy Optimization
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.23095