Historian: Reducing Manual Validation in APR Benchmarking via Evidence-Based Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moslemi, Sahand, Lami, Mayasah, Koyuncu, Anil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMShot: Reducing snapshot testing maintenance via LLMs
von: Kaynak, Ergün Batuhan, et al.
Veröffentlicht: (2025)
von: Kaynak, Ergün Batuhan, et al.
Veröffentlicht: (2025)
Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks
von: Bosnak, Emir, et al.
Veröffentlicht: (2025)
von: Bosnak, Emir, et al.
Veröffentlicht: (2025)
Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
von: Pehlivan, Fatih, et al.
Veröffentlicht: (2025)
von: Pehlivan, Fatih, et al.
Veröffentlicht: (2025)
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
von: Ma, Yinghang, et al.
Veröffentlicht: (2025)
von: Ma, Yinghang, et al.
Veröffentlicht: (2025)
Beyond Localization: Recoverable Headroom and Residual Frontier in Repository-Level RAG-APR
von: Zhao, Pengtao, et al.
Veröffentlicht: (2026)
von: Zhao, Pengtao, et al.
Veröffentlicht: (2026)
Human-Aligned Code Readability Assessment with Large Language Models
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
T5APR: Empowering Automated Program Repair across Languages through Checkpoint Ensemble
von: Gharibi, Reza, et al.
Veröffentlicht: (2023)
von: Gharibi, Reza, et al.
Veröffentlicht: (2023)
Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
von: Li, Yuanhao, et al.
Veröffentlicht: (2026)
von: Li, Yuanhao, et al.
Veröffentlicht: (2026)
Test smells in LLM-Generated Unit Tests
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2024)
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2024)
Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2025)
Large-scale, Independent and Comprehensive study of the power of LLMs for test case generation
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2024)
von: Ouédraogo, Wendkûuni C., et al.
Veröffentlicht: (2024)
Evidence Tetris in the Pixelated World of Validity Threats
von: Wyrich, Marvin, et al.
Veröffentlicht: (2024)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2024)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
von: Wang, Longtian, et al.
Veröffentlicht: (2025)
von: Wang, Longtian, et al.
Veröffentlicht: (2025)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
Studying Quality Improvements Recommended via Manual and Automated Code Review
von: Crupi, Giuseppe, et al.
Veröffentlicht: (2026)
von: Crupi, Giuseppe, et al.
Veröffentlicht: (2026)
Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
von: Dai, Yihan, et al.
Veröffentlicht: (2025)
von: Dai, Yihan, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Bug Reproduction for Android Apps via Pre-Assessment of Visual Effects
von: Xiao, Xiangyang, et al.
Veröffentlicht: (2026)
von: Xiao, Xiangyang, et al.
Veröffentlicht: (2026)
Validity-Preserving Delta Debugging via Generator Trace Reduction
von: Ren, Luyao, et al.
Veröffentlicht: (2024)
von: Ren, Luyao, et al.
Veröffentlicht: (2024)
Single-Language Evidence Is Insufficient for Automated Logging: A Multilingual Benchmark and Empirical Study with LLMs
von: Zhong, Renyi, et al.
Veröffentlicht: (2026)
von: Zhong, Renyi, et al.
Veröffentlicht: (2026)
Towards Evidence-Based Tech Hiring Pipelines
von: Brown, Chris, et al.
Veröffentlicht: (2025)
von: Brown, Chris, et al.
Veröffentlicht: (2025)
Ghost Echoes Revealed: Benchmarking Maintainability Metrics and Machine Learning Predictions Against Human Assessments
von: Borg, Markus, et al.
Veröffentlicht: (2024)
von: Borg, Markus, et al.
Veröffentlicht: (2024)
Accelerating Patch Validation for Program Repair with Interception-Based Execution Scheduling
von: Xiao, Yuan-An, et al.
Veröffentlicht: (2023)
von: Xiao, Yuan-An, et al.
Veröffentlicht: (2023)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
von: Zhang, Zejun, et al.
Veröffentlicht: (2026)
von: Zhang, Zejun, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of Manually Created Equivalent Mutants
von: Straubinger, Philipp, et al.
Veröffentlicht: (2024)
von: Straubinger, Philipp, et al.
Veröffentlicht: (2024)
Guideline for Manual Process Discovery in Industrial IoT
von: Kölbel, Linda, et al.
Veröffentlicht: (2024)
von: Kölbel, Linda, et al.
Veröffentlicht: (2024)
Reducing Alert Fatigue via AI-Assisted Negotiation: A Case for Dependabot
von: Kula, Raula Gaikovina
Veröffentlicht: (2025)
von: Kula, Raula Gaikovina
Veröffentlicht: (2025)
Data-Driven Evidence-Based Syntactic Sugar Design
von: OBrien, David, et al.
Veröffentlicht: (2024)
von: OBrien, David, et al.
Veröffentlicht: (2024)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
An Empirical Study of Reducing AV1 Decoder Complexity and Energy Consumption via Encoder Parameter Tuning
von: Vibhoothi, Vibhoothi, et al.
Veröffentlicht: (2025)
von: Vibhoothi, Vibhoothi, et al.
Veröffentlicht: (2025)
Exploring the Evidence-Based SE Beliefs of Generative AI Tools
von: Brown, Chris, et al.
Veröffentlicht: (2024)
von: Brown, Chris, et al.
Veröffentlicht: (2024)
RUM: Rule+LLM-Based Comprehensive Assessment on Testing Skills
von: Wang, Yue, et al.
Veröffentlicht: (2025)
von: Wang, Yue, et al.
Veröffentlicht: (2025)
Leveraging Language Models to Discover Evidence-Based Actions for OSS Sustainability
von: Khan, Nafiz Imtiaz, et al.
Veröffentlicht: (2026)
von: Khan, Nafiz Imtiaz, et al.
Veröffentlicht: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Generative AI in Evidence-Based Software Engineering: A White Paper
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
von: Esposito, Matteo, et al.
Veröffentlicht: (2024)
Validity in Design Science
von: Larsen, K., et al.
Veröffentlicht: (2025)
von: Larsen, K., et al.
Veröffentlicht: (2025)
Metamorphic Testing for Smart Contract Validation:A Case Study of Ethereum-Based Crowdfunding Contracts
von: Villanueva, Irving Jared, et al.
Veröffentlicht: (2025)
von: Villanueva, Irving Jared, et al.
Veröffentlicht: (2025)
LogSage: An LLM-Based Framework for CI/CD Failure Detection and Remediation with Industrial Validation
von: Xu, Weiyuan, et al.
Veröffentlicht: (2025)
von: Xu, Weiyuan, et al.
Veröffentlicht: (2025)
Reusing Model Validation Methods for the Continuous Validation of Digital Twins of Cyber-Physical Systems
von: Mertens, Joost, et al.
Veröffentlicht: (2025)
von: Mertens, Joost, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLMShot: Reducing snapshot testing maintenance via LLMs
von: Kaynak, Ergün Batuhan, et al.
Veröffentlicht: (2025) -
Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks
von: Bosnak, Emir, et al.
Veröffentlicht: (2025) -
Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
von: Pehlivan, Fatih, et al.
Veröffentlicht: (2025) -
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
von: Ma, Yinghang, et al.
Veröffentlicht: (2025) -
Beyond Localization: Recoverable Headroom and Residual Frontier in Repository-Level RAG-APR
von: Zhao, Pengtao, et al.
Veröffentlicht: (2026)