PIPer: On-Device Environment Setup via Online Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kovrigin, Alexander, Eliseeva, Aleksandra, Grotov, Konstantin, Bogomolov, Egor, Zharov, Yaroslav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024)
by: Kovrigin, Alexander, et al.
Published: (2024)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026)
by: Slinko, Igor, et al.
Published: (2026)
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)
by: Bogomolov, Egor, et al.
Published: (2024)
Untangling Knots: Leveraging LLM for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
Themisto: Jupyter-Based Runtime Benchmark
by: Grotov, Konstantin, et al.
Published: (2025)
by: Grotov, Konstantin, et al.
Published: (2025)
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
by: Tsvetkov, Petr, et al.
Published: (2024)
by: Tsvetkov, Petr, et al.
Published: (2024)
Challenge on Optimization of Context Collection for Code Completion
by: Ustalov, Dmitry, et al.
Published: (2025)
by: Ustalov, Dmitry, et al.
Published: (2025)
Dynamic Retrieval-Augmented Generation
by: Shapkin, Anton, et al.
Published: (2023)
by: Shapkin, Anton, et al.
Published: (2023)
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
by: Glukhov, Evgeniy, et al.
Published: (2025)
by: Glukhov, Evgeniy, et al.
Published: (2025)
Stack Trace Deduplication: Faster, More Accurately, and in More Realistic Scenarios
by: Shibaev, Egor, et al.
Published: (2024)
by: Shibaev, Egor, et al.
Published: (2024)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
by: Arora, Avi, et al.
Published: (2025)
by: Arora, Avi, et al.
Published: (2025)
Tool-Augmented LLMs as a Universal Interface for IDEs
by: Zharov, Yaroslav, et al.
Published: (2024)
by: Zharov, Yaroslav, et al.
Published: (2024)
Debug Smarter, Not Harder: AI Agents for Error Resolution in Computational Notebooks
by: Grotov, Konstantin, et al.
Published: (2024)
by: Grotov, Konstantin, et al.
Published: (2024)
MIST-RL: Mutation-based Incremental Suite Testing via Reinforcement Learning
by: Zhu, Sicheng, et al.
Published: (2026)
by: Zhu, Sicheng, et al.
Published: (2026)
Reinforcement Learning for Online Testing of Autonomous Driving Systems: a Replication and Extension Study
by: Giamattei, Luca, et al.
Published: (2024)
by: Giamattei, Luca, et al.
Published: (2024)
Hammer: Robust Function-Calling for On-Device Language Models via Function Masking
by: Lin, Qiqiang, et al.
Published: (2024)
by: Lin, Qiqiang, et al.
Published: (2024)
Usability and Performance Analysis of Embedded Development Environment for On-device Learning
by: Scaffi, Enzo, et al.
Published: (2024)
by: Scaffi, Enzo, et al.
Published: (2024)
A Reference Architecture of Reinforcement Learning Frameworks
by: Liu, Xiaoran, et al.
Published: (2026)
by: Liu, Xiaoran, et al.
Published: (2026)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
by: Deshpande, Darshan, et al.
Published: (2026)
by: Deshpande, Darshan, et al.
Published: (2026)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
by: Bouchoucha, Rached, et al.
Published: (2024)
by: Bouchoucha, Rached, et al.
Published: (2024)
Complex Model Transformations by Reinforcement Learning with Uncertain Human Guidance
by: Dagenais, Kyanna, et al.
Published: (2025)
by: Dagenais, Kyanna, et al.
Published: (2025)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
by: Li, Xin-Ye, et al.
Published: (2026)
by: Li, Xin-Ye, et al.
Published: (2026)
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
by: Zolfagharian, Amirhossein, et al.
Published: (2023)
by: Zolfagharian, Amirhossein, et al.
Published: (2023)
Recommending Pre-Trained Models for IoT Devices
by: Patil, Parth V., et al.
Published: (2024)
by: Patil, Parth V., et al.
Published: (2024)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
by: Yuan, Danlong, et al.
Published: (2026)
by: Yuan, Danlong, et al.
Published: (2026)
ADReFT: Adaptive Decision Repair for Safe Autonomous Driving via Reinforcement Fine-Tuning
by: Cheng, Mingfei, et al.
Published: (2025)
by: Cheng, Mingfei, et al.
Published: (2025)
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
Automatic Generation of High-Performance RL Environments
by: Karten, Seth, et al.
Published: (2026)
by: Karten, Seth, et al.
Published: (2026)
The Impact of Environment Configurations on the Stability of AI-Enabled Systems
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
Assuring the Safety of Reinforcement Learning Components: AMLAS-RL
by: Imrie, Calum Corrie, et al.
Published: (2025)
by: Imrie, Calum Corrie, et al.
Published: (2025)
Mellum: Production-Grade in-IDE Contextual Code Completion with Multi-File Project Understanding
by: Pavlichenko, Nikita, et al.
Published: (2025)
by: Pavlichenko, Nikita, et al.
Published: (2025)
Teaching an Online Multi-Institutional Research Level Software Engineering Course with Industry -- an Experience Report
by: Jalote, Pankaj, et al.
Published: (2025)
by: Jalote, Pankaj, et al.
Published: (2025)
Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-Learning
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
BatCoder: Self-Supervised Bidirectional Code-Documentation Learning via Back-Translation
by: Xu, Jingwen, et al.
Published: (2026)
by: Xu, Jingwen, et al.
Published: (2026)
Similar Items
-
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025) -
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
by: Kovrigin, Alexander, et al.
Published: (2024) -
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026) -
GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git
by: Lindenbauer, Tobias, et al.
Published: (2025) -
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
by: Bogomolov, Egor, et al.
Published: (2024)