On Randomness in Agentic Evals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bjarnason, Bjarni Haukur, Silva, André, Monperrus, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Supersonic: Learning to Generate Source Code Optimizations in C/C++
von: Chen, Zimin, et al.
Veröffentlicht: (2023)
von: Chen, Zimin, et al.
Veröffentlicht: (2023)
RepairBench: Leaderboard of Frontier Models for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2024)
von: Silva, André, et al.
Veröffentlicht: (2024)
Generative AI to Generate Test Data Generators
von: Baudry, Benoit, et al.
Veröffentlicht: (2024)
von: Baudry, Benoit, et al.
Veröffentlicht: (2024)
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2023)
von: Silva, André, et al.
Veröffentlicht: (2023)
Bootstrapping Coding Agents: The Specification Is the Program
von: Monperrus, Martin
Veröffentlicht: (2026)
von: Monperrus, Martin
Veröffentlicht: (2026)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
von: Xu, Weiwei, et al.
Veröffentlicht: (2024)
Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces
von: Silva, André, et al.
Veröffentlicht: (2025)
von: Silva, André, et al.
Veröffentlicht: (2025)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts
von: Andersson, Vivi, et al.
Veröffentlicht: (2025)
von: Andersson, Vivi, et al.
Veröffentlicht: (2025)
ITER: Iterative Neural Repair for Multi-Location Patches
von: Ye, He, et al.
Veröffentlicht: (2023)
von: Ye, He, et al.
Veröffentlicht: (2023)
GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries
von: Hou, Shuyang, et al.
Veröffentlicht: (2025)
von: Hou, Shuyang, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
von: Storhaug, André, et al.
Veröffentlicht: (2024)
von: Storhaug, André, et al.
Veröffentlicht: (2024)
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
von: Charoenwet, Wachiraphan, et al.
Veröffentlicht: (2026)
von: Charoenwet, Wachiraphan, et al.
Veröffentlicht: (2026)
Scaling Test-Time Compute for Agentic Coding
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
Leveraging XP and CRISP-DM for Agile Data Science Projects
von: Shimaoka, Andre Massahiro, et al.
Veröffentlicht: (2025)
von: Shimaoka, Andre Massahiro, et al.
Veröffentlicht: (2025)
CP-Agent: Agentic Constraint Programming
von: Szeider, Stefan
Veröffentlicht: (2025)
von: Szeider, Stefan
Veröffentlicht: (2025)
AFlow: Automating Agentic Workflow Generation
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
Are Sparse Autoencoders Useful for Java Function Bug Detection?
von: Melo, Rui, et al.
Veröffentlicht: (2025)
von: Melo, Rui, et al.
Veröffentlicht: (2025)
Learning to Compose for Cross-domain Agentic Workflow Generation
von: Wang, Jialiang, et al.
Veröffentlicht: (2026)
von: Wang, Jialiang, et al.
Veröffentlicht: (2026)
Assessing Large Language Models for Automated Feedback Generation in Learning Programming Problem Solving
von: Silva, Priscylla, et al.
Veröffentlicht: (2025)
von: Silva, Priscylla, et al.
Veröffentlicht: (2025)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
CigaR: Cost-efficient Program Repair with LLMs
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
The Unreasonable Effectiveness of Open Science in AI: A Replication Study
von: Gundersen, Odd Erik, et al.
Veröffentlicht: (2024)
von: Gundersen, Odd Erik, et al.
Veröffentlicht: (2024)
Adaptive Detection of Software Aging under Workload Shift
von: Silva, Rafael Jose Moura, et al.
Veröffentlicht: (2025)
von: Silva, Rafael Jose Moura, et al.
Veröffentlicht: (2025)
David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design?
von: Shankar, Shashwat, et al.
Veröffentlicht: (2025)
von: Shankar, Shashwat, et al.
Veröffentlicht: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
von: Andrews, Martin, et al.
Veröffentlicht: (2025)
von: Andrews, Martin, et al.
Veröffentlicht: (2025)
Adaptation of XAI to Auto-tuning for Numerical Libraries
von: Aoki, Shota, et al.
Veröffentlicht: (2024)
von: Aoki, Shota, et al.
Veröffentlicht: (2024)
RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2025)
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
von: Xu, WeiZhe, et al.
Veröffentlicht: (2026)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
R-LAM: Reproducibility-Constrained Large Action Models for Scientific Workflow Automation
von: Sureshkumar, Suriya
Veröffentlicht: (2026)
von: Sureshkumar, Suriya
Veröffentlicht: (2026)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
von: Yu, Shasha, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
Standing on the Shoulders of Giants: Stabilized Knowledge Distillation for Cross--Language Code Clone Detection
von: Khajezade, Mohamad, et al.
Veröffentlicht: (2026)
von: Khajezade, Mohamad, et al.
Veröffentlicht: (2026)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Supersonic: Learning to Generate Source Code Optimizations in C/C++
von: Chen, Zimin, et al.
Veröffentlicht: (2023) -
RepairBench: Leaderboard of Frontier Models for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2024) -
Generative AI to Generate Test Data Generators
von: Baudry, Benoit, et al.
Veröffentlicht: (2024) -
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2023) -
Bootstrapping Coding Agents: The Specification Is the Program
von: Monperrus, Martin
Veröffentlicht: (2026)