A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tabassum, Anika, Hossain, Md Sifat, Arefin, Md. Fahim, Islam, Tariqul, Zaman, Tarannum Shaila
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916022778331136
author Tabassum, Anika
Hossain, Md Sifat
Arefin, Md. Fahim
Islam, Tariqul
Zaman, Tarannum Shaila
author_facet Tabassum, Anika
Hossain, Md Sifat
Arefin, Md. Fahim
Islam, Tariqul
Zaman, Tarannum Shaila
contents Large Language Models (LLMs) demonstrate strong potential for automated code generation, yet their ability to iteratively refine solutions using execution feedback remains underexplored. Competitive programming offers an ideal testbed for this investigation, as it demands end-to-end algorithmic reasoning, precise implementation under strict computational constraints, and complete functional correctness with rigorous evaluation. In this paper, we present A-ProS, an autonomous AI agent that solves competitive programming problems through a hybrid multi-model feedback framework separating solution generation from specialized debugging. A-ProS combines ChatGPT-based generators (GPT-4 and GPT-5) with three debugging critics: Codestral-2508, Llama-3.3-70B, and DeepSeek-R1, under a 2 x 3 factorial design. We evaluate six workflows on 367 problems from ICPC World Finals (2011-2024) and Codeforces (rated 1200-1800). The results show that GPT-5 workflows improve from 39 initial accepted solutions to 85-90 after three refinement rounds, while GPT-4 improves from 15 to 31-38. A controlled ablation on 47 problems shows that stateful refinement outperforms stateless approaches by 8.5-10.6 percentage points and reduces repeated failures by up to 3.5x. Compared to baseline agent loops, A-ProS achieves over 2x greater gains, highlighting the importance of persistent context and multi-model feedback for reliable autonomous program synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18073
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
Tabassum, Anika
Hossain, Md Sifat
Arefin, Md. Fahim
Islam, Tariqul
Zaman, Tarannum Shaila
Software Engineering
Artificial Intelligence
Large Language Models (LLMs) demonstrate strong potential for automated code generation, yet their ability to iteratively refine solutions using execution feedback remains underexplored. Competitive programming offers an ideal testbed for this investigation, as it demands end-to-end algorithmic reasoning, precise implementation under strict computational constraints, and complete functional correctness with rigorous evaluation. In this paper, we present A-ProS, an autonomous AI agent that solves competitive programming problems through a hybrid multi-model feedback framework separating solution generation from specialized debugging. A-ProS combines ChatGPT-based generators (GPT-4 and GPT-5) with three debugging critics: Codestral-2508, Llama-3.3-70B, and DeepSeek-R1, under a 2 x 3 factorial design. We evaluate six workflows on 367 problems from ICPC World Finals (2011-2024) and Codeforces (rated 1200-1800). The results show that GPT-5 workflows improve from 39 initial accepted solutions to 85-90 after three refinement rounds, while GPT-4 improves from 15 to 31-38. A controlled ablation on 47 problems shows that stateful refinement outperforms stateless approaches by 8.5-10.6 percentage points and reduces repeated failures by up to 3.5x. Compared to baseline agent loops, A-ProS achieves over 2x greater gains, highlighting the importance of persistent context and multi-model feedback for reliable autonomous program synthesis.
title A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2605.18073