Beyond pip install: Evaluating LLM Agents for the Automated Installation of Python Projects
Fuente:
arXiv
Salvato in:
| Autori principali: | Milliken, Louis, Kang, Sungmin, Yoo, Shin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Identifying Inaccurate Descriptions in LLM-generated Code Comments via Test Execution
di: Kang, Sungmin, et al.
Pubblicazione: (2024)
di: Kang, Sungmin, et al.
Pubblicazione: (2024)
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
di: Kang, Sungmin, et al.
Pubblicazione: (2023)
di: Kang, Sungmin, et al.
Pubblicazione: (2023)
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
di: Kim, Naryeong, et al.
Pubblicazione: (2024)
di: Kim, Naryeong, et al.
Pubblicazione: (2024)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
di: Cho, Hyunjoon, et al.
Pubblicazione: (2025)
di: Cho, Hyunjoon, et al.
Pubblicazione: (2025)
Predictive Prompt Analysis
di: Lee, Jae Yong, et al.
Pubblicazione: (2025)
di: Lee, Jae Yong, et al.
Pubblicazione: (2025)
Finding the Needle in the Crash Stack: Industrial-Scale Crash Root Cause Localization with AutoCrashFL
di: Kang, Sungmin, et al.
Pubblicazione: (2025)
di: Kang, Sungmin, et al.
Pubblicazione: (2025)
Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap
di: Kim, Naryeong, et al.
Pubblicazione: (2026)
di: Kim, Naryeong, et al.
Pubblicazione: (2026)
Evaluating LLM Agents on Automated Software Analysis Tasks
di: Bouzenia, Islem, et al.
Pubblicazione: (2026)
di: Bouzenia, Islem, et al.
Pubblicazione: (2026)
METAMON: Finding Inconsistencies between Program Documentation and Behavior using Metamorphic LLM Queries
di: Lee, Hyeonseok, et al.
Pubblicazione: (2025)
di: Lee, Hyeonseok, et al.
Pubblicazione: (2025)
Real-World Fault Detection for C-Extended Python Projects with Automated Unit Test Generation
di: Berg, Lucas, et al.
Pubblicazione: (2026)
di: Berg, Lucas, et al.
Pubblicazione: (2026)
MigMate: A VS Code Extension for LLM-based Library Migration of Python Projects
di: Kebede, Matthias, et al.
Pubblicazione: (2026)
di: Kebede, Matthias, et al.
Pubblicazione: (2026)
LLM Agents for Automated Dependency Upgrades
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
AutoCodeSherpa: Symbolic Explanations in AI Coding Agents
di: Kang, Sungmin, et al.
Pubblicazione: (2025)
di: Kang, Sungmin, et al.
Pubblicazione: (2025)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
di: Gajjar, Jugal, et al.
Pubblicazione: (2025)
di: Gajjar, Jugal, et al.
Pubblicazione: (2025)
Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
di: Xiang, Jiahong, et al.
Pubblicazione: (2026)
di: Xiang, Jiahong, et al.
Pubblicazione: (2026)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
di: Lin, Yixi, et al.
Pubblicazione: (2026)
di: Lin, Yixi, et al.
Pubblicazione: (2026)
Adapting Installation Instructions in Rapidly Evolving Software Ecosystems
di: Gao, Haoyu, et al.
Pubblicazione: (2023)
di: Gao, Haoyu, et al.
Pubblicazione: (2023)
Leveraging LLM Agents for Automated Video Game Testing
di: Wang, Chengjia, et al.
Pubblicazione: (2025)
di: Wang, Chengjia, et al.
Pubblicazione: (2025)
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
di: Saxena, Divyanshu, et al.
Pubblicazione: (2025)
Beyond Rules: LLM-Powered Linting for Quantum Programs
di: Cassieri, Pietro, et al.
Pubblicazione: (2026)
di: Cassieri, Pietro, et al.
Pubblicazione: (2026)
PCART: Automated Repair of Python API Parameter Compatibility Issues
di: Zhang, Shuai, et al.
Pubblicazione: (2024)
di: Zhang, Shuai, et al.
Pubblicazione: (2024)
Combining Type Inference and Automated Unit Test Generation for Python
di: Krodinger, Lukas, et al.
Pubblicazione: (2025)
di: Krodinger, Lukas, et al.
Pubblicazione: (2025)
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
Towards Automated Governance: A DSL for Human-Agent Collaboration in Software Projects
di: Ait, Adem, et al.
Pubblicazione: (2025)
di: Ait, Adem, et al.
Pubblicazione: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
di: Zhang, Ke, et al.
Pubblicazione: (2025)
di: Zhang, Ke, et al.
Pubblicazione: (2025)
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
di: Islam, Niful, et al.
Pubblicazione: (2026)
di: Islam, Niful, et al.
Pubblicazione: (2026)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
di: Abbassi, Altaf Allah, et al.
Pubblicazione: (2025)
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
Comparing ML-Specific and General Python Code Smells Across Project Characteristics
di: Agh, Halimeh, et al.
Pubblicazione: (2026)
di: Agh, Halimeh, et al.
Pubblicazione: (2026)
PyGress: Tool for Analyzing the Progression of Code Proficiency in Python OSS Projects
di: Charatvaraphan, Rujiphart, et al.
Pubblicazione: (2025)
di: Charatvaraphan, Rujiphart, et al.
Pubblicazione: (2025)
Performance Smells in ML and Non-ML Python Projects: A Comparative Study
di: Belias, François, et al.
Pubblicazione: (2025)
di: Belias, François, et al.
Pubblicazione: (2025)
Automated Refactoring of Non-Idiomatic Python Code: A Differentiated Replication with LLMs
di: Midolo, Alessandro, et al.
Pubblicazione: (2025)
di: Midolo, Alessandro, et al.
Pubblicazione: (2025)
PCREQ: Automated Inference of Compatible Requirements for Python Third-party Library Upgrades
di: Lei, Huashan, et al.
Pubblicazione: (2025)
di: Lei, Huashan, et al.
Pubblicazione: (2025)
AgentFL: Scaling LLM-based Fault Localization to Project-Level Context
di: Qin, Yihao, et al.
Pubblicazione: (2024)
di: Qin, Yihao, et al.
Pubblicazione: (2024)
Open Source Software Development Tool Installation: Challenges and Strategies For Novice Developers
di: Salerno, Larissa, et al.
Pubblicazione: (2024)
di: Salerno, Larissa, et al.
Pubblicazione: (2024)
A Task-Level Evaluation of AI Agents in Open-Source Projects
di: Rahman, Shojibur, et al.
Pubblicazione: (2026)
di: Rahman, Shojibur, et al.
Pubblicazione: (2026)
Testing in the Evolving World of DL Systems:Insights from Python GitHub Projects
di: Ali, Qurban, et al.
Pubblicazione: (2024)
di: Ali, Qurban, et al.
Pubblicazione: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Identifying Inaccurate Descriptions in LLM-generated Code Comments via Test Execution
di: Kang, Sungmin, et al.
Pubblicazione: (2024) -
A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
di: Kang, Sungmin, et al.
Pubblicazione: (2023) -
Lachesis: Predicting LLM Inference Accuracy using Structural Properties of Reasoning Paths
di: Kim, Naryeong, et al.
Pubblicazione: (2024) -
COSMosFL: Ensemble of Small Language Models for Fault Localisation
di: Cho, Hyunjoon, et al.
Pubblicazione: (2025) -
Predictive Prompt Analysis
di: Lee, Jae Yong, et al.
Pubblicazione: (2025)