SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Boxi, Cao, Yang, Zhang, Yuzhong, Lin, Liting, Xu, Junjielong, Zhong, Zhiqing, Xu, Qinghua, Wang, Guancheng, Cao, Jialun, Cheung, Shing-Chi, He, Pinjia, Briand, Lionel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
di: Yu, Boxi, et al.
Pubblicazione: (2026)
di: Yu, Boxi, et al.
Pubblicazione: (2026)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
di: Yu, Boxi, et al.
Pubblicazione: (2025)
di: Yu, Boxi, et al.
Pubblicazione: (2025)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
di: Li, Meiziniu, et al.
Pubblicazione: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
di: Li, Meiziniu, et al.
Pubblicazione: (2022)
LLM-Driven Cost-Effective Requirements Change Impact Analysis
di: Etezadi, Romina, et al.
Pubblicazione: (2025)
di: Etezadi, Romina, et al.
Pubblicazione: (2025)
Technical Debt Management: The Road Ahead for Successful Software Delivery
di: Avgeriou, Paris, et al.
Pubblicazione: (2024)
di: Avgeriou, Paris, et al.
Pubblicazione: (2024)
Tests4Py: A Benchmark for System Testing
di: Smytzek, Marius, et al.
Pubblicazione: (2023)
di: Smytzek, Marius, et al.
Pubblicazione: (2023)
Talk is Cheap, Logic is Hard: Benchmarking LLMs on Post-Condition Formalization
di: Prasetya, I. S. W. B., et al.
Pubblicazione: (2026)
di: Prasetya, I. S. W. B., et al.
Pubblicazione: (2026)
AgentOps: Enabling Observability of LLM Agents
di: Dong, Liming, et al.
Pubblicazione: (2024)
di: Dong, Liming, et al.
Pubblicazione: (2024)
SIADAFIX: issue description response for adaptive program repair
di: Cao, Xin, et al.
Pubblicazione: (2025)
di: Cao, Xin, et al.
Pubblicazione: (2025)
IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development
di: Moreira, Jasmine
Pubblicazione: (2026)
di: Moreira, Jasmine
Pubblicazione: (2026)
CONGRA: Benchmarking Automatic Conflict Resolution
di: Zhang, Qingyu, et al.
Pubblicazione: (2024)
di: Zhang, Qingyu, et al.
Pubblicazione: (2024)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
di: Kessel, Marcus
Pubblicazione: (2024)
di: Kessel, Marcus
Pubblicazione: (2024)
PyPackIT: Automated Research Software Engineering for Scientific Python Applications on GitHub
di: Ariamajd, Armin, et al.
Pubblicazione: (2025)
di: Ariamajd, Armin, et al.
Pubblicazione: (2025)
Architectural Patterns for Designing Quantum Artificial Intelligence Systems
di: Klymenko, Mykhailo, et al.
Pubblicazione: (2024)
di: Klymenko, Mykhailo, et al.
Pubblicazione: (2024)
Comprehensive Evaluation of Large Language Models on Software Engineering Tasks: A Multi-Task Benchmark
di: Gunawan, Go Frendi, et al.
Pubblicazione: (2026)
di: Gunawan, Go Frendi, et al.
Pubblicazione: (2026)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
di: Li, Meiziniu, et al.
Pubblicazione: (2026)
di: Li, Meiziniu, et al.
Pubblicazione: (2026)
Scalable and Secure AI Inference in Healthcare: A Comparative Benchmarking of FastAPI and Triton Inference Server on Kubernetes
di: Ali, Ratul
Pubblicazione: (2026)
di: Ali, Ratul
Pubblicazione: (2026)
Factors that Contribute to the Success of a Software Organisation's DevOps Environment: A Systematic Review
di: Gwangwadza, Ashley, et al.
Pubblicazione: (2022)
di: Gwangwadza, Ashley, et al.
Pubblicazione: (2022)
Code Documentation and Analysis to Secure Software Development
di: Attie, Paul, et al.
Pubblicazione: (2024)
di: Attie, Paul, et al.
Pubblicazione: (2024)
QUT: A Unit Testing Framework for Quantum Subroutines
di: Klymenko, Mykhailo V., et al.
Pubblicazione: (2025)
di: Klymenko, Mykhailo V., et al.
Pubblicazione: (2025)
GBM Returns the Best Prediction Performance among Regression Approaches: A Case Study of Stack Overflow Code Quality
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
Proof of Concept as a First-Class Architectural Decision Instrument
di: Antognolli, Bruno Fernando, et al.
Pubblicazione: (2026)
di: Antognolli, Bruno Fernando, et al.
Pubblicazione: (2026)
A History Equivalence Algorithm for Dynamic Process Migration
di: Bakshi, Gargi, et al.
Pubblicazione: (2024)
di: Bakshi, Gargi, et al.
Pubblicazione: (2024)
Early-Stage Requirements Transformation Approaches: A Systematic Review
di: Letsholo, Keletso J.
Pubblicazione: (2024)
di: Letsholo, Keletso J.
Pubblicazione: (2024)
AddressWatcher: Sanitizer-Based Localization of Memory Leak Fixes
di: Murali, Aniruddhan, et al.
Pubblicazione: (2024)
di: Murali, Aniruddhan, et al.
Pubblicazione: (2024)
CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
di: Li, Shiyang, et al.
Pubblicazione: (2026)
di: Li, Shiyang, et al.
Pubblicazione: (2026)
Model Generation with LLMs: From Requirements to UML Sequence Diagrams
di: Ferrari, Alessio, et al.
Pubblicazione: (2024)
di: Ferrari, Alessio, et al.
Pubblicazione: (2024)
Code Less to Code More: Streamlining Language Server Protocol and Type System Development for Language Families
di: Bruzzone, Federico, et al.
Pubblicazione: (2025)
di: Bruzzone, Federico, et al.
Pubblicazione: (2025)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
di: Dinu, Ion George, et al.
Pubblicazione: (2026)
Addressing Visual Impairments with Model-Driven Engineering: A Systematic Literature Review
di: Michael, Judith, et al.
Pubblicazione: (2025)
di: Michael, Judith, et al.
Pubblicazione: (2025)
Behavior Trees and State Machines in Robotics Applications
di: Ghzouli, Razan, et al.
Pubblicazione: (2022)
di: Ghzouli, Razan, et al.
Pubblicazione: (2022)
The Kieker Observability Framework Version 2
di: Yang, Shinhyung, et al.
Pubblicazione: (2025)
di: Yang, Shinhyung, et al.
Pubblicazione: (2025)
Variability Modeling of Products, Processes, and Resources in Cyber-Physical Production Systems Engineering
di: Meixner, Kristof, et al.
Pubblicazione: (2024)
di: Meixner, Kristof, et al.
Pubblicazione: (2024)
Validating API Design Requirements for Interoperability: A Static Analysis Approach Using OpenAPI
di: Sundberg, Edwin, et al.
Pubblicazione: (2025)
di: Sundberg, Edwin, et al.
Pubblicazione: (2025)
Validating Formal Specifications with LLM-generated Test Cases
di: Cunha, Alcino, et al.
Pubblicazione: (2025)
di: Cunha, Alcino, et al.
Pubblicazione: (2025)
GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
di: Heilman, Alex, et al.
Pubblicazione: (2026)
di: Heilman, Alex, et al.
Pubblicazione: (2026)
Synthesizing Test Cases for Narrowing Specification Candidates
di: Cunha, Alcino, et al.
Pubblicazione: (2025)
di: Cunha, Alcino, et al.
Pubblicazione: (2025)
Comparing Human and LLM Generated Code: The Jury is Still Out!
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
di: Licorish, Sherlock A., et al.
Pubblicazione: (2025)
ContextBench: A Benchmark for Context Retrieval in Coding Agents
di: Li, Han, et al.
Pubblicazione: (2026)
di: Li, Han, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
di: Yu, Boxi, et al.
Pubblicazione: (2026) -
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
di: Yu, Boxi, et al.
Pubblicazione: (2025) -
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
di: Li, Meiziniu, et al.
Pubblicazione: (2024) -
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
di: Li, Meiziniu, et al.
Pubblicazione: (2022) -
LLM-Driven Cost-Effective Requirements Change Impact Analysis
di: Etezadi, Romina, et al.
Pubblicazione: (2025)