QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Doyeon, Lyou, Eunyi, Cho, Hyunsoo, Kim, Sookyung, Lee, Joonseok, Choi, Jaemoo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911426946269184
author Lee, Doyeon
Lyou, Eunyi
Cho, Hyunsoo
Kim, Sookyung
Lee, Joonseok
Choi, Jaemoo
author_facet Lee, Doyeon
Lyou, Eunyi
Cho, Hyunsoo
Kim, Sookyung
Lee, Joonseok
Choi, Jaemoo
contents GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04620
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
Lee, Doyeon
Lyou, Eunyi
Cho, Hyunsoo
Kim, Sookyung
Lee, Joonseok
Choi, Jaemoo
Machine Learning
GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training.
title QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
topic Machine Learning
url https://arxiv.org/abs/2602.04620