QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911426946269184 |
|---|---|
| author | Lee, Doyeon Lyou, Eunyi Cho, Hyunsoo Kim, Sookyung Lee, Joonseok Choi, Jaemoo |
| author_facet | Lee, Doyeon Lyou, Eunyi Cho, Hyunsoo Kim, Sookyung Lee, Joonseok Choi, Jaemoo |
| contents | GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_04620 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning Lee, Doyeon Lyou, Eunyi Cho, Hyunsoo Kim, Sookyung Lee, Joonseok Choi, Jaemoo Machine Learning GRPO-style reinforcement learning (RL)-based LLM fine-tuning algorithms have recently gained popularity. Relying on heuristic trust-region approximations, however, they can lead to brittle optimization behavior, as global importance-ratio clipping and group-wise normalization fail to regulate samples whose importance ratios fall outside the clipping range. We propose Query-Adaptive Trust-Region policy Optimization (QUATRO), which directly enforces trust-region constraints through a principled optimization. This yields a clear and interpretable objective that enables explicit control over policy updates and stable, entropy-controlled optimization, with a stabilizer terms arising intrinsically from the exact trust-region formulation. Empirically verified on diverse mathematical reasoning benchmarks, QUATRO shows stable training under increased policy staleness and aggressive learning rates, maintaining well-controlled entropy throughout training. |
| title | QUATRO: Query-Adaptive Trust Region Policy Optimization for LLM Fine-tuning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2602.04620 |