Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiazheng, Fu, Ziche, Shen, Junrui, Zhao, Yunbin, Zhang, Yunke, Xi, Zhiheng, Ma, Long, An, Chenxin, Zhang, Zhihao, Liu, Shichun, Zhu, Dingwei, Dou, Shihan, Liu, Shaofan, Li, Han, Zhou, Wiggin, Adams, Aiden, Gui, Tao, Huang, Fei, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910220876251136
author Zhang, Jiazheng
Fu, Ziche
Shen, Junrui
Zhao, Yunbin
Zhang, Yunke
Xi, Zhiheng
Ma, Long
An, Chenxin
Zhang, Zhihao
Liu, Shichun
Zhu, Dingwei
Dou, Shihan
Liu, Shaofan
Li, Han
Zhou, Wiggin
Adams, Aiden
Gui, Tao
Huang, Fei
Zhang, Qi
Huang, Xuanjing
author_facet Zhang, Jiazheng
Fu, Ziche
Shen, Junrui
Zhao, Yunbin
Zhang, Yunke
Xi, Zhiheng
Ma, Long
An, Chenxin
Zhang, Zhihao
Liu, Shichun
Zhu, Dingwei
Dou, Shihan
Liu, Shaofan
Li, Han
Zhou, Wiggin
Adams, Aiden
Gui, Tao
Huang, Fei
Zhang, Qi
Huang, Xuanjing
contents Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existing entropy-aware methods mainly regulate entropy through global objectives, while the token-level mechanism by which sampled policy updates reshape policy entropy remains underexplored. In this work, we develop a theoretical framework of entropy mechanics in RLVR. Our analysis yields a first-order approximation of the entropy change, giving rise to entropy polarity, a signed token-level quantity that predicts how much a sampled update expands or contracts entropy. This analysis further reveals a structural asymmetry: reinforcing frequent high-probability tokens triggers contraction tendencies, whereas expansive tendencies typically require lower-probability samples or stronger distributional correction. Empirically, we show that entropy polarity reliably predicts entropy changes, and that positive and negative polarity branches play complementary roles in preserving exploration while strengthening exploitation. Building on these insights, we propose Polarity-Aware Policy Optimization (PAPO), which preserves both polarity branches and implements entropy control through advantage reweighting. With the empirical entropy trajectory as an online phase signal, PAPO adaptively reallocates optimization pressure between entropy-expanding and entropy-contracting updates. Experiments on mathematical reasoning and agentic benchmarks show that PAPO consistently outperforms competitive baselines, while delivering superior training efficiency and substantial reward improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11775
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
Zhang, Jiazheng
Fu, Ziche
Shen, Junrui
Zhao, Yunbin
Zhang, Yunke
Xi, Zhiheng
Ma, Long
An, Chenxin
Zhang, Zhihao
Liu, Shichun
Zhu, Dingwei
Dou, Shihan
Liu, Shaofan
Li, Han
Zhou, Wiggin
Adams, Aiden
Gui, Tao
Huang, Fei
Zhang, Qi
Huang, Xuanjing
Machine Learning
Computation and Language
Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existing entropy-aware methods mainly regulate entropy through global objectives, while the token-level mechanism by which sampled policy updates reshape policy entropy remains underexplored. In this work, we develop a theoretical framework of entropy mechanics in RLVR. Our analysis yields a first-order approximation of the entropy change, giving rise to entropy polarity, a signed token-level quantity that predicts how much a sampled update expands or contracts entropy. This analysis further reveals a structural asymmetry: reinforcing frequent high-probability tokens triggers contraction tendencies, whereas expansive tendencies typically require lower-probability samples or stronger distributional correction. Empirically, we show that entropy polarity reliably predicts entropy changes, and that positive and negative polarity branches play complementary roles in preserving exploration while strengthening exploitation. Building on these insights, we propose Polarity-Aware Policy Optimization (PAPO), which preserves both polarity branches and implements entropy control through advantage reweighting. With the empirical entropy trajectory as an online phase signal, PAPO adaptively reallocates optimization pressure between entropy-expanding and entropy-contracting updates. Experiments on mathematical reasoning and agentic benchmarks show that PAPO consistently outperforms competitive baselines, while delivering superior training efficiency and substantial reward improvements.
title Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.11775