DyPyBench: A Benchmark of Executable Python Software
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bouzenia, Islem, Krishan, Bajaj Piyush, Pradel, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
von: Joos, Pascal, et al.
Veröffentlicht: (2025)
von: Joos, Pascal, et al.
Veröffentlicht: (2025)
Issue2Test: Generating Reproducing Test Cases from Issue Reports
von: Nashid, Noor, et al.
Veröffentlicht: (2025)
von: Nashid, Noor, et al.
Veröffentlicht: (2025)
PyTy: Repairing Static Type Errors in Python
von: Chow, Yiu Wai, et al.
Veröffentlicht: (2024)
von: Chow, Yiu Wai, et al.
Veröffentlicht: (2024)
Treefix: Enabling Execution with a Tree of Prefixes
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
Names Are All You Need: Effective and Safe Regression Test Selection for Python
von: Wang, You, et al.
Veröffentlicht: (2026)
von: Wang, You, et al.
Veröffentlicht: (2026)
From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
von: Zhu, Hao-Nan, et al.
Veröffentlicht: (2025)
von: Zhu, Hao-Nan, et al.
Veröffentlicht: (2025)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
AgentStepper: Interactive Debugging of Software Development Agents
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
Testora: Using Natural Language Intent to Detect Behavioral Regressions
von: Pradel, Michael
Veröffentlicht: (2025)
von: Pradel, Michael
Veröffentlicht: (2025)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
FauxPy: A Fault Localization Tool for Python
von: Rezaalipour, Mohammad, et al.
Veröffentlicht: (2024)
von: Rezaalipour, Mohammad, et al.
Veröffentlicht: (2024)
HQPEF-Py: Metrics, Python Patterns, and Guidance for Evaluating Hybrid Quantum Programs
von: Osei, Michael Adjei, et al.
Veröffentlicht: (2025)
von: Osei, Michael Adjei, et al.
Veröffentlicht: (2025)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
Artisan: Agentic Artifact Evaluation
von: Baek, Doehyun, et al.
Veröffentlicht: (2026)
von: Baek, Doehyun, et al.
Veröffentlicht: (2026)
Agentic AI Software Engineers: Programming with Trust
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
Less is More? An Empirical Study on Configuration Issues in Python PyPI Ecosystem
von: Peng, Yun, et al.
Veröffentlicht: (2023)
von: Peng, Yun, et al.
Veröffentlicht: (2023)
PyTrim: A Practical Tool for Reducing Python Dependency Bloat
von: Karakatsanis, Konstantinos, et al.
Veröffentlicht: (2025)
von: Karakatsanis, Konstantinos, et al.
Veröffentlicht: (2025)
PyTracer: Automatically profiling numerical instabilities in Python
von: Chatelain, Yohan, et al.
Veröffentlicht: (2021)
von: Chatelain, Yohan, et al.
Veröffentlicht: (2021)
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
Execution-Aware Program Reduction for WebAssembly via Record and Replay
von: Baek, Doehyun, et al.
Veröffentlicht: (2025)
von: Baek, Doehyun, et al.
Veröffentlicht: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
PyExamine A Comprehensive, UnOpinionated Smell Detection Tool for Python
von: Shivashankar, Karthik, et al.
Veröffentlicht: (2025)
von: Shivashankar, Karthik, et al.
Veröffentlicht: (2025)
HaPy-Bug -- Human Annotated Python Bug Resolution Dataset
von: Przymus, Piotr, et al.
Veröffentlicht: (2025)
von: Przymus, Piotr, et al.
Veröffentlicht: (2025)
PyPulse: A Python Library for Biosignal Imputation
von: Gao, Kevin, et al.
Veröffentlicht: (2024)
von: Gao, Kevin, et al.
Veröffentlicht: (2024)
TypeEvalPy: A Micro-benchmarking Framework for Python Type Inference Tools
von: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Veröffentlicht: (2023)
von: Venkatesh, Ashwin Prasad Shivarpatna, et al.
Veröffentlicht: (2023)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
RippleGUItester: Change-Aware Exploratory Testing
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
PyGress: Tool for Analyzing the Progression of Code Proficiency in Python OSS Projects
von: Charatvaraphan, Rujiphart, et al.
Veröffentlicht: (2025)
von: Charatvaraphan, Rujiphart, et al.
Veröffentlicht: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
Gender Disparities in Contributions, Leadership, and Collaboration: An Exploratory Study on Software Systems Research
von: Cynthia, Shamse Tasnim, et al.
Veröffentlicht: (2024)
von: Cynthia, Shamse Tasnim, et al.
Veröffentlicht: (2024)
SafePyScript: A Web-Based Solution for Machine Learning-Driven Vulnerability Detection in Python
von: Farasat, Talaya, et al.
Veröffentlicht: (2024)
von: Farasat, Talaya, et al.
Veröffentlicht: (2024)
BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024) -
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025) -
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026) -
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024) -
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
von: Joos, Pascal, et al.
Veröffentlicht: (2025)