Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shrivastava, Vaishnavi, Awadallah, Ahmed, Balachandran, Vidhisha, Garg, Shivam, Behl, Harkirat, Papailiopoulos, Dimitris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909735828062208
author Shrivastava, Vaishnavi
Awadallah, Ahmed
Balachandran, Vidhisha
Garg, Shivam
Behl, Harkirat
Papailiopoulos, Dimitris
author_facet Shrivastava, Vaishnavi
Awadallah, Ahmed
Balachandran, Vidhisha
Garg, Shivam
Behl, Harkirat
Papailiopoulos, Dimitris
contents Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length--inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many tokens are merely "filler": repetitive, verbose text that makes no real progress. We introduce GFPO (Group Filtered Policy Optimization), which curbs this length explosion by sampling larger groups per problem during training and filtering responses to train on based on two key metrics: (1) response length and (2) token efficiency: reward per token ratio. By sampling more at training time, we teach models to think less at inference time. On the Phi-4-reasoning model, GFPO cuts GRPO's length inflation by 46-71% across challenging STEM and coding benchmarks (AIME 24/25, GPQA, Omni-MATH, LiveCodeBench) while maintaining accuracy. Optimizing for reward per token further increases reductions in length inflation to 71-85%. We also propose Adaptive Difficulty GFPO, which dynamically allocates more training resources to harder problems based on real-time difficulty estimates, improving the balance between computational efficiency and accuracy especially on difficult questions. GFPO demonstrates that increased training-time compute directly translates to reduced test-time compute--a simple yet effective trade-off for efficient reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Shrivastava, Vaishnavi
Awadallah, Ahmed
Balachandran, Vidhisha
Garg, Shivam
Behl, Harkirat
Papailiopoulos, Dimitris
Computation and Language
Machine Learning
Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length--inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many tokens are merely "filler": repetitive, verbose text that makes no real progress. We introduce GFPO (Group Filtered Policy Optimization), which curbs this length explosion by sampling larger groups per problem during training and filtering responses to train on based on two key metrics: (1) response length and (2) token efficiency: reward per token ratio. By sampling more at training time, we teach models to think less at inference time. On the Phi-4-reasoning model, GFPO cuts GRPO's length inflation by 46-71% across challenging STEM and coding benchmarks (AIME 24/25, GPQA, Omni-MATH, LiveCodeBench) while maintaining accuracy. Optimizing for reward per token further increases reductions in length inflation to 71-85%. We also propose Adaptive Difficulty GFPO, which dynamically allocates more training resources to harder problems based on real-time difficulty estimates, improving the balance between computational efficiency and accuracy especially on difficult questions. GFPO demonstrates that increased training-time compute directly translates to reduced test-time compute--a simple yet effective trade-off for efficient reasoning.
title Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.09726