GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yikun, Wang, Yibin, Wang, Dianyi, Peng, Zimian, Guo, Qipeng, Tao, Dacheng, Wang, Jiaqi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908978815959040
author Wang, Yikun
Wang, Yibin
Wang, Dianyi
Peng, Zimian
Guo, Qipeng
Tao, Dacheng
Wang, Jiaqi
author_facet Wang, Yikun
Wang, Yibin
Wang, Dianyi
Peng, Zimian
Guo, Qipeng
Tao, Dacheng
Wang, Jiaqi
contents Recent progress in large language models (LLMs) has boosted mathematical reasoning, yet geometry remains challenging where auxiliary construction is often essential. Prior methods either underperform or depend on very large models (e.g., GPT-4o), making them costly. We argue that reinforcement learning with verifiable rewards (e.g., GRPO) can train smaller models to couple auxiliary construction with solid geometric reasoning. However, naively applying GRPO yields unconditional rewards, encouraging indiscriminate and sometimes harmful constructions. We propose Group Contrastive Policy Optimization (GCPO), an RL framework with two components: (1) Group Contrastive Masking, which assigns positive/negative construction rewards based on contextual utility, and (2) a Length Reward that encourages longer reasoning chains. On top of GCPO, we build GeometryZero, an affordable family of geometry reasoning models that selectively use auxiliary construction. Experiments on Geometry3K and MathVista show GeometryZero consistently outperforms RL baselines (e.g., GRPO, ToRL). The code has been available at https://github.com/ekonwang/GeometryZero.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization
Wang, Yikun
Wang, Yibin
Wang, Dianyi
Peng, Zimian
Guo, Qipeng
Tao, Dacheng
Wang, Jiaqi
Computation and Language
Recent progress in large language models (LLMs) has boosted mathematical reasoning, yet geometry remains challenging where auxiliary construction is often essential. Prior methods either underperform or depend on very large models (e.g., GPT-4o), making them costly. We argue that reinforcement learning with verifiable rewards (e.g., GRPO) can train smaller models to couple auxiliary construction with solid geometric reasoning. However, naively applying GRPO yields unconditional rewards, encouraging indiscriminate and sometimes harmful constructions. We propose Group Contrastive Policy Optimization (GCPO), an RL framework with two components: (1) Group Contrastive Masking, which assigns positive/negative construction rewards based on contextual utility, and (2) a Length Reward that encourages longer reasoning chains. On top of GCPO, we build GeometryZero, an affordable family of geometry reasoning models that selectively use auxiliary construction. Experiments on Geometry3K and MathVista show GeometryZero consistently outperforms RL baselines (e.g., GRPO, ToRL). The code has been available at https://github.com/ekonwang/GeometryZero.
title GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization
topic Computation and Language
url https://arxiv.org/abs/2506.07160