IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yinhan, Zhu, Yaochen, Shi, Mingjia, Zheng, Wendy, Su, Lin, Wang, Xiaoqing, Guo, Qi, Li, Jundong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916063021629440
author He, Yinhan
Zhu, Yaochen
Shi, Mingjia
Zheng, Wendy
Su, Lin
Wang, Xiaoqing
Guo, Qi
Li, Jundong
author_facet He, Yinhan
Zhu, Yaochen
Shi, Mingjia
Zheng, Wendy
Su, Lin
Wang, Xiaoqing
Guo, Qi
Li, Jundong
contents Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We revisit token-efficient post-training and argue that existing sequence-level reward-shaping methods offer limited control over how reasoning effort is allocated across tokens. To bridge the gap, we propose IAPO, an information-theoretic post-training framework that assigns token-wise advantages based on each token's conditional mutual information (MI) with the final answer. This yields an explicit, principled mechanism for identifying informative reasoning steps and suppressing low-utility exploration. We provide a theoretical analysis showing that our IAPO can induce monotonic reductions in reasoning verbosity without harming correctness. Empirically, IAPO consistently improves reasoning accuracy while reducing reasoning length by up to 36%, outperforming existing token-efficient RL methods across various reasoning datasets. Extensive empirical evaluations demonstrate that information-aware advantage shaping is a powerful and general direction for token-efficient post-training. The code is available at https://github.com/YinhanHe123/IAPO.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19049
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
He, Yinhan
Zhu, Yaochen
Shi, Mingjia
Zheng, Wendy
Su, Lin
Wang, Xiaoqing
Guo, Qi
Li, Jundong
Computation and Language
Machine Learning
Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We revisit token-efficient post-training and argue that existing sequence-level reward-shaping methods offer limited control over how reasoning effort is allocated across tokens. To bridge the gap, we propose IAPO, an information-theoretic post-training framework that assigns token-wise advantages based on each token's conditional mutual information (MI) with the final answer. This yields an explicit, principled mechanism for identifying informative reasoning steps and suppressing low-utility exploration. We provide a theoretical analysis showing that our IAPO can induce monotonic reductions in reasoning verbosity without harming correctness. Empirically, IAPO consistently improves reasoning accuracy while reducing reasoning length by up to 36%, outperforming existing token-efficient RL methods across various reasoning datasets. Extensive empirical evaluations demonstrate that information-aware advantage shaping is a powerful and general direction for token-efficient post-training. The code is available at https://github.com/YinhanHe123/IAPO.
title IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.19049