PIPer: On-Device Environment Setup via Online Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kovrigin, Alexander, Eliseeva, Aleksandra, Grotov, Konstantin, Bogomolov, Egor, Zharov, Yaroslav
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914094839234560
author Kovrigin, Alexander
Eliseeva, Aleksandra
Grotov, Konstantin
Bogomolov, Egor
Zharov, Yaroslav
author_facet Kovrigin, Alexander
Eliseeva, Aleksandra
Grotov, Konstantin
Bogomolov, Egor
Zharov, Yaroslav
contents Environment setup-the process of configuring the system to work with a specific software project-represents a persistent challenge in Software Engineering (SE). Automated environment setup methods could assist developers by providing fully configured environments for arbitrary repositories without manual effort. This also helps SE researchers to scale execution-based benchmarks. However, recent studies reveal that even state-of-the-art Large Language Models (LLMs) achieve limited success in automating this task. To address this limitation, we tune a specialized model for environment setup. We combine supervised fine-tuning for generating correct Bash scripts and Reinforcement Learning with Verifiable Rewards (RLVR) to adapt it to the task of environment setup. On EnvBench-Python, our method enables Qwen3-8B (a model runnable on consumer hardware) to perform on par with larger models-Qwen3-32B and GPT-4o. The training code and model checkpoints are available online: https://github.com/JetBrains-Research/PIPer.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25455
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIPer: On-Device Environment Setup via Online Reinforcement Learning
Kovrigin, Alexander
Eliseeva, Aleksandra
Grotov, Konstantin
Bogomolov, Egor
Zharov, Yaroslav
Machine Learning
Artificial Intelligence
Software Engineering
Environment setup-the process of configuring the system to work with a specific software project-represents a persistent challenge in Software Engineering (SE). Automated environment setup methods could assist developers by providing fully configured environments for arbitrary repositories without manual effort. This also helps SE researchers to scale execution-based benchmarks. However, recent studies reveal that even state-of-the-art Large Language Models (LLMs) achieve limited success in automating this task. To address this limitation, we tune a specialized model for environment setup. We combine supervised fine-tuning for generating correct Bash scripts and Reinforcement Learning with Verifiable Rewards (RLVR) to adapt it to the task of environment setup. On EnvBench-Python, our method enables Qwen3-8B (a model runnable on consumer hardware) to perform on par with larger models-Qwen3-32B and GPT-4o. The training code and model checkpoints are available online: https://github.com/JetBrains-Research/PIPer.
title PIPer: On-Device Environment Setup via Online Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2509.25455