Addressing Data Leakage in HumanEval Using Combinatorial Test Design
Fuente:
arXiv
Saved in:
| Main Authors: | Bradbury, Jeremy S., More, Riddhi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
by: Ahmed, Sheikh Nazib, et al.
Published: (2026)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026)
by: Terragni, Valerio
Published: (2026)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
by: Chang, Hung-Fu, et al.
Published: (2025)
by: Chang, Hung-Fu, et al.
Published: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
by: Jana, Prithwish, et al.
Published: (2023)
by: Jana, Prithwish, et al.
Published: (2023)
Automated structural testing of LLM-based agents: methods, framework, and case studies
by: Kohl, Jens, et al.
Published: (2026)
by: Kohl, Jens, et al.
Published: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Automating Domain-Driven Design: Experience with a Prompting Framework
by: Eisenreich, Tobias, et al.
Published: (2026)
by: Eisenreich, Tobias, et al.
Published: (2026)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
by: Ashrafi, Nazmus
Published: (2026)
by: Ashrafi, Nazmus
Published: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
by: Baldonado, Juan Manuel, et al.
Published: (2025)
by: Baldonado, Juan Manuel, et al.
Published: (2025)
EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
by: Costa, Igor, et al.
Published: (2025)
by: Costa, Igor, et al.
Published: (2025)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
by: Kessel, Marcus
Published: (2024)
by: Kessel, Marcus
Published: (2024)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
by: Eisenreich, Tobias, et al.
Published: (2025)
by: Eisenreich, Tobias, et al.
Published: (2025)
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
by: Trooskens, Geert, et al.
Published: (2026)
by: Trooskens, Geert, et al.
Published: (2026)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
by: Holt, Samuel, et al.
Published: (2023)
by: Holt, Samuel, et al.
Published: (2023)
OODEval: Evaluating Large Language Models on Object-Oriented Design
by: Xiao, Bingxu, et al.
Published: (2026)
by: Xiao, Bingxu, et al.
Published: (2026)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
by: Dinu, Ion George, et al.
Published: (2026)
by: Dinu, Ion George, et al.
Published: (2026)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
From Scientific Texts to Verifiable Code: Automating the Process with Transformers
by: Wang, Changjie, et al.
Published: (2025)
by: Wang, Changjie, et al.
Published: (2025)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
by: Miller, Jan
Published: (2026)
by: Miller, Jan
Published: (2026)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
by: Ranasinghe, Nishath Rajiv, et al.
Published: (2025)
by: Ranasinghe, Nishath Rajiv, et al.
Published: (2025)
N-Version Assessment and Enhancement of Generative AI
by: Kessel, Marcus, et al.
Published: (2024)
by: Kessel, Marcus, et al.
Published: (2024)
Morescient GAI for Software Engineering (Extended Version)
by: Kessel, Marcus, et al.
Published: (2024)
by: Kessel, Marcus, et al.
Published: (2024)
Similar Items
-
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025) -
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025) -
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025) -
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
by: Ahmed, Sheikh Nazib, et al.
Published: (2026) -
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)