FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
Fuente:
arXiv
Salvato in:
| Autori principali: | May, Victor, Misra, Diganta, Luo, Yanqi, Sridhar, Anjali, Gehring, Justine, Junior, Silvio Soares Ribeiro |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
di: Misra, Diganta, et al.
Pubblicazione: (2025)
di: Misra, Diganta, et al.
Pubblicazione: (2025)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
di: Islah, Nizar, et al.
Pubblicazione: (2024)
di: Islah, Nizar, et al.
Pubblicazione: (2024)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
di: Amin, Nishil, et al.
Pubblicazione: (2026)
di: Amin, Nishil, et al.
Pubblicazione: (2026)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
di: Liu, Linbo, et al.
Pubblicazione: (2025)
di: Liu, Linbo, et al.
Pubblicazione: (2025)
ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java
di: Pavuluri, Advait, et al.
Pubblicazione: (2026)
di: Pavuluri, Advait, et al.
Pubblicazione: (2026)
Quality Evaluation of COBOL to Java Code Transformation
di: Froimovich, Shmulik, et al.
Pubblicazione: (2025)
di: Froimovich, Shmulik, et al.
Pubblicazione: (2025)
jscefr: A Framework to Evaluate the Code Proficiency for JavaScript
di: Ragkhitwetsagul, Chaiyong, et al.
Pubblicazione: (2024)
di: Ragkhitwetsagul, Chaiyong, et al.
Pubblicazione: (2024)
MigMate: A VS Code Extension for LLM-based Library Migration of Python Projects
di: Kebede, Matthias, et al.
Pubblicazione: (2026)
di: Kebede, Matthias, et al.
Pubblicazione: (2026)
Acknowledging Good Java Code with Code Perfumes
di: Straubinger, Philipp, et al.
Pubblicazione: (2024)
di: Straubinger, Philipp, et al.
Pubblicazione: (2024)
Generating Java Methods: An Empirical Assessment of Four AI-Based Code Assistants
di: Corso, Vincenzo, et al.
Pubblicazione: (2024)
di: Corso, Vincenzo, et al.
Pubblicazione: (2024)
GitBug-Java: A Reproducible Benchmark of Recent Java Bugs
di: Silva, André, et al.
Pubblicazione: (2024)
di: Silva, André, et al.
Pubblicazione: (2024)
Serializing Java Objects in Plain Code
di: Wachter, Julian, et al.
Pubblicazione: (2024)
di: Wachter, Julian, et al.
Pubblicazione: (2024)
Decoding the Configuration of AI Coding Agents: Insights from Claude Code Projects
di: Santos, Helio Victor F., et al.
Pubblicazione: (2025)
di: Santos, Helio Victor F., et al.
Pubblicazione: (2025)
Environment-in-the-Loop: Rethinking Code Migration with LLM-based Agents
di: Li, Xiang, et al.
Pubblicazione: (2026)
di: Li, Xiang, et al.
Pubblicazione: (2026)
Detection of Technical Debt in Java Source Code
di: Hai, Nam Le, et al.
Pubblicazione: (2024)
di: Hai, Nam Le, et al.
Pubblicazione: (2024)
Demystifying and Assessing Code Understandability in Java Decompilation
di: Qin, Ruixin, et al.
Pubblicazione: (2024)
di: Qin, Ruixin, et al.
Pubblicazione: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
Resolving Java Code Repository Issues with iSWE Agent
di: Ganhotra, Jatin, et al.
Pubblicazione: (2026)
di: Ganhotra, Jatin, et al.
Pubblicazione: (2026)
Zero-shot Evaluation of Deep Learning for Java Code Clone Detection
di: Heinze, Thomas S.
Pubblicazione: (2026)
di: Heinze, Thomas S.
Pubblicazione: (2026)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
di: Cao, Jialun, et al.
Pubblicazione: (2024)
di: Cao, Jialun, et al.
Pubblicazione: (2024)
MapReplay: Trace-Driven Benchmark Generation for Java HashMap
di: Schiavio, Filippo, et al.
Pubblicazione: (2026)
di: Schiavio, Filippo, et al.
Pubblicazione: (2026)
Code Review Automation Via Multi-task Federated LLM -- An Empirical Study
di: Kumar, Jahnavi, et al.
Pubblicazione: (2024)
di: Kumar, Jahnavi, et al.
Pubblicazione: (2024)
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
di: Xie, Bang, et al.
Pubblicazione: (2026)
di: Xie, Bang, et al.
Pubblicazione: (2026)
Do Comments and Expertise Still Matter? An Experiment on Programmers' Adoption of AI-Generated JavaScript Code
di: Li, Changwen, et al.
Pubblicazione: (2025)
di: Li, Changwen, et al.
Pubblicazione: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
di: Zhao, Songwen, et al.
Pubblicazione: (2025)
di: Zhao, Songwen, et al.
Pubblicazione: (2025)
Code Review Agent Benchmark
di: Zhang, Yuntong, et al.
Pubblicazione: (2026)
di: Zhang, Yuntong, et al.
Pubblicazione: (2026)
A Soundness and Precision Benchmark for Java Debloating Tools
di: Klauke, Jonas, et al.
Pubblicazione: (2025)
di: Klauke, Jonas, et al.
Pubblicazione: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
di: Li, Jialin, et al.
Pubblicazione: (2026)
di: Li, Jialin, et al.
Pubblicazione: (2026)
Scalable Thread-Safety Analysis of Java Classes with CodeQL
di: Jåtten, Bjørnar Haugstad, et al.
Pubblicazione: (2025)
di: Jåtten, Bjørnar Haugstad, et al.
Pubblicazione: (2025)
VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
di: Han, Liutong, et al.
Pubblicazione: (2025)
di: Han, Liutong, et al.
Pubblicazione: (2025)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
di: Sonwane, Atharv, et al.
Pubblicazione: (2026)
di: Sonwane, Atharv, et al.
Pubblicazione: (2026)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
di: Guo, Chengquan, et al.
Pubblicazione: (2024)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
Code Reborn AI-Driven Legacy Systems Modernization from COBOL to Java
di: Bandarupalli, Gopichand
Pubblicazione: (2025)
di: Bandarupalli, Gopichand
Pubblicazione: (2025)
Fingerprinting AI Coding Agents on GitHub
di: Ghaleb, Taher A.
Pubblicazione: (2026)
di: Ghaleb, Taher A.
Pubblicazione: (2026)
Formal Methods Meets Readability: Auto-Documenting JML Java Code
di: Abad, Juan Carlos Recio, et al.
Pubblicazione: (2025)
di: Abad, Juan Carlos Recio, et al.
Pubblicazione: (2025)
Generating Accurate OpenAPI Descriptions from Java Source Code
di: Lercher, Alexander, et al.
Pubblicazione: (2024)
di: Lercher, Alexander, et al.
Pubblicazione: (2024)
Automated Refactoring of Legacy JavaScript Code to ES6 Modules
di: Paltoglou, Katerina, et al.
Pubblicazione: (2021)
di: Paltoglou, Katerina, et al.
Pubblicazione: (2021)
Unmasking the Genuine Type Inference Capabilities of LLMs for Java Code Snippets
di: Dong, Yiwen, et al.
Pubblicazione: (2025)
di: Dong, Yiwen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
di: Misra, Diganta, et al.
Pubblicazione: (2025) -
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
di: Islah, Nizar, et al.
Pubblicazione: (2024) -
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
di: Amin, Nishil, et al.
Pubblicazione: (2026) -
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
di: Liu, Linbo, et al.
Pubblicazione: (2025) -
ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java
di: Pavuluri, Advait, et al.
Pubblicazione: (2026)