Once Upon an Input: Reasoning via Per-Instance Program Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stein, Adam, Velingker, Neelay, Naik, Mayur, Wong, Eric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908614133809152
author Stein, Adam
Velingker, Neelay
Naik, Mayur
Wong, Eric
author_facet Stein, Adam
Velingker, Neelay
Naik, Mayur
Wong, Eric
contents Large language models (LLMs) excel at zero-shot inference but continue to struggle with complex, multi-step reasoning. Recent methods that augment LLMs with intermediate reasoning steps such as Chain of Thought (CoT) and Program of Thought (PoT) improve performance but often produce undesirable solutions, especially in algorithmic domains. We introduce Per-Instance Program Synthesis (PIPS), a method that generates and refines programs at the instance-level using structural feedback without relying on task-specific guidance or explicit test cases. To further improve performance, PIPS incorporates a confidence metric that dynamically chooses between direct inference and program synthesis on a per-instance basis. Experiments across three frontier LLMs and 30 benchmarks including all tasks of Big Bench Extra Hard (BBEH), visual question answering tasks, relational reasoning tasks, and mathematical reasoning tasks show that PIPS improves the absolute harmonic mean accuracy by up to 8.6% and 9.4% compared to PoT and CoT respectively, and reduces undesirable program generations by 65.1% on the algorithmic tasks compared to PoT with Gemini-2.0-Flash.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Once Upon an Input: Reasoning via Per-Instance Program Synthesis
Stein, Adam
Velingker, Neelay
Naik, Mayur
Wong, Eric
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) excel at zero-shot inference but continue to struggle with complex, multi-step reasoning. Recent methods that augment LLMs with intermediate reasoning steps such as Chain of Thought (CoT) and Program of Thought (PoT) improve performance but often produce undesirable solutions, especially in algorithmic domains. We introduce Per-Instance Program Synthesis (PIPS), a method that generates and refines programs at the instance-level using structural feedback without relying on task-specific guidance or explicit test cases. To further improve performance, PIPS incorporates a confidence metric that dynamically chooses between direct inference and program synthesis on a per-instance basis. Experiments across three frontier LLMs and 30 benchmarks including all tasks of Big Bench Extra Hard (BBEH), visual question answering tasks, relational reasoning tasks, and mathematical reasoning tasks show that PIPS improves the absolute harmonic mean accuracy by up to 8.6% and 9.4% compared to PoT and CoT respectively, and reduces undesirable program generations by 65.1% on the algorithmic tasks compared to PoT with Gemini-2.0-Flash.
title Once Upon an Input: Reasoning via Per-Instance Program Synthesis
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.22849