Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yoon, Youngsik, Lee, Sungjae, Song, Seockbean, Wang, Siwei, Chen, Wei, Ok, Jungseul
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2605.07248
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911661564100608
author Yoon, Youngsik
Lee, Sungjae
Song, Seockbean
Wang, Siwei
Chen, Wei
Ok, Jungseul
author_facet Yoon, Youngsik
Lee, Sungjae
Song, Seockbean
Wang, Siwei
Chen, Wei
Ok, Jungseul
contents Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Models (LLMs). However, most existing methods adopt a rigid Planning-before-Trial (PbT) policy, which inefficiently allocates test-time compute by incurring planning overhead even on directly solvable problems. We propose Planning-after-Trial (PaT), an adaptive policy for code generation that invokes a planner only upon verification failure. This adaptive policy naturally enables a heterogeneous model configuration: a cost-efficient model handles generation attempts, while a powerful model is reserved for targeted planning interventions. Empirically, across multiple benchmarks and model families, our approach significantly advances the cost-performance Pareto frontier. Notably, our heterogeneous configuration achieves performance comparable to a large homogeneous model while reducing inference cost by approximately 69\%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaT: Planning-after-Trial for Efficient Test-Time Code Generation
Yoon, Youngsik
Lee, Sungjae
Song, Seockbean
Wang, Siwei
Chen, Wei
Ok, Jungseul
Computation and Language
Machine Learning
Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Models (LLMs). However, most existing methods adopt a rigid Planning-before-Trial (PbT) policy, which inefficiently allocates test-time compute by incurring planning overhead even on directly solvable problems. We propose Planning-after-Trial (PaT), an adaptive policy for code generation that invokes a planner only upon verification failure. This adaptive policy naturally enables a heterogeneous model configuration: a cost-efficient model handles generation attempts, while a powerful model is reserved for targeted planning interventions. Empirically, across multiple benchmarks and model families, our approach significantly advances the cost-performance Pareto frontier. Notably, our heterogeneous configuration achieves performance comparable to a large homogeneous model while reducing inference cost by approximately 69\%.
title PaT: Planning-after-Trial for Efficient Test-Time Code Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.07248