GIPO: Gaussian Importance Sampling Policy Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lu, Chengxuan, Zhang, Zhenquan, Wang, Shukuan, Lin, Qunzhi, Sun, Baigui, Liu, Yang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918368998588416
author Lu, Chengxuan
Zhang, Zhenquan
Wang, Shukuan
Lin, Qunzhi
Sun, Baigui
Liu, Yang
author_facet Lu, Chengxuan
Zhang, Zhenquan
Wang, Shukuan
Lin, Qunzhi
Sun, Baigui
Liu, Yang
contents Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this challenge, GIPO (Gaussian Importance sampling Policy Optimization) is proposed as a policy optimization objective based on truncated importance sampling, replacing hard clipping with a log-ratio-based Gaussian trust weight to softly damp extreme importance ratios while maintaining non-zero gradients. Theoretical analysis shows that GIPO introduces an implicit, tunable constraint on the update magnitude, while concentration bounds guarantee robustness and stability under finite-sample estimation. Experimental results show that GIPO achieves state-of-the-art performance among clipping-based baselines across a wide range of replay buffer sizes, from near on-policy to highly stale data, while exhibiting superior bias--variance trade-off, high training stability and improved sample efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03955
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GIPO: Gaussian Importance Sampling Policy Optimization
Lu, Chengxuan
Zhang, Zhenquan
Wang, Shukuan
Lin, Qunzhi
Sun, Baigui
Liu, Yang
Machine Learning
Artificial Intelligence
Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this challenge, GIPO (Gaussian Importance sampling Policy Optimization) is proposed as a policy optimization objective based on truncated importance sampling, replacing hard clipping with a log-ratio-based Gaussian trust weight to softly damp extreme importance ratios while maintaining non-zero gradients. Theoretical analysis shows that GIPO introduces an implicit, tunable constraint on the update magnitude, while concentration bounds guarantee robustness and stability under finite-sample estimation. Experimental results show that GIPO achieves state-of-the-art performance among clipping-based baselines across a wide range of replay buffer sizes, from near on-policy to highly stale data, while exhibiting superior bias--variance trade-off, high training stability and improved sample efficiency.
title GIPO: Gaussian Importance Sampling Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.03955