AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Zihan, Wang, Xiaohan, Yang, Hexiong, Chai, Jiajun, Cao, Jie, Yin, Guojun, Lin, Wei, He, Ran
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911376567435264
author Lin, Zihan
Wang, Xiaohan
Yang, Hexiong
Chai, Jiajun
Cao, Jie
Yin, Guojun
Lin, Wei
He, Ran
author_facet Lin, Zihan
Wang, Xiaohan
Yang, Hexiong
Chai, Jiajun
Cao, Jie
Yin, Guojun
Lin, Wei
He, Ran
contents While Reinforcement Learning (RL) shows promise in training tool-use Large Language Models (LLMs) using verifiable outcome rewards, existing methods largely overlook the potential of reasoning rewards based on chain-of-thought quality for better tool utilization. Furthermore, naïvely combining reasoning and outcome rewards may yield suboptimal performance or conflict with the primary optimization objective. To address this, we propose Advantage-Weighted Policy Optimization (AWPO), a principled RL framework that adaptively integrates reasoning rewards into advantage estimation to improve tool-use performance. AWPO incorporates variance-aware gating and difficulty-aware weighting to adaptively modulate advantages from reasoning signals based on group-relative statistics, alongside a tailored clipping mechanism for stable optimization. Extensive experiments demonstrate that AWPO achieves state-of-the-art performance across standard tool-use benchmarks, significantly outperforming strong baselines and leading closed-source models in challenging multi-turn scenarios. Notably, with exceptional parameter efficiency, our 4B model surpasses Grok-4 by $16.0\%$ in multi-turn accuracy while preserving generalization capability on the out-of-distribution MMLU-Pro benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
Lin, Zihan
Wang, Xiaohan
Yang, Hexiong
Chai, Jiajun
Cao, Jie
Yin, Guojun
Lin, Wei
He, Ran
Computation and Language
While Reinforcement Learning (RL) shows promise in training tool-use Large Language Models (LLMs) using verifiable outcome rewards, existing methods largely overlook the potential of reasoning rewards based on chain-of-thought quality for better tool utilization. Furthermore, naïvely combining reasoning and outcome rewards may yield suboptimal performance or conflict with the primary optimization objective. To address this, we propose Advantage-Weighted Policy Optimization (AWPO), a principled RL framework that adaptively integrates reasoning rewards into advantage estimation to improve tool-use performance. AWPO incorporates variance-aware gating and difficulty-aware weighting to adaptively modulate advantages from reasoning signals based on group-relative statistics, alongside a tailored clipping mechanism for stable optimization. Extensive experiments demonstrate that AWPO achieves state-of-the-art performance across standard tool-use benchmarks, significantly outperforming strong baselines and leading closed-source models in challenging multi-turn scenarios. Notably, with exceptional parameter efficiency, our 4B model surpasses Grok-4 by $16.0\%$ in multi-turn accuracy while preserving generalization capability on the out-of-distribution MMLU-Pro benchmark.
title AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
topic Computation and Language
url https://arxiv.org/abs/2512.19126