G-Zero: Self-Play for Open-Ended Generation from Zero Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chengsong, Liu, Haolin, Zheng, Tong, Dai, Runpeng, Huang, Langlin, Li, Jinyuan, Li, Zongxia, Wei, Zhepei, Meng, Yu, Huang, Jiaxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911670374236160
author Huang, Chengsong
Liu, Haolin
Zheng, Tong
Dai, Runpeng
Huang, Langlin
Li, Jinyuan
Li, Zongxia
Wei, Zhepei
Meng, Yu
Huang, Jiaxin
author_facet Huang, Chengsong
Liu, Haolin
Zheng, Tong
Dai, Runpeng
Huang, Langlin
Li, Jinyuan
Li, Zongxia
Wei, Zhepei
Meng, Yu
Huang, Jiaxin
contents Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-$δ$, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09959
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle G-Zero: Self-Play for Open-Ended Generation from Zero Data
Huang, Chengsong
Liu, Haolin
Zheng, Tong
Dai, Runpeng
Huang, Langlin
Li, Jinyuan
Li, Zongxia
Wei, Zhepei
Meng, Yu
Huang, Jiaxin
Machine Learning
Artificial Intelligence
Computation and Language
Emerging Technologies
Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-$δ$, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.
title G-Zero: Self-Play for Open-Ended Generation from Zero Data
topic Machine Learning
Artificial Intelligence
Computation and Language
Emerging Technologies
url https://arxiv.org/abs/2605.09959