JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Jialun, Chen, Zhiyong, Wu, Jiarong, Cheung, Shing-chi, Xu, Chang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
Can Large Language Models Model Programs Formally?
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
von: Chen, Songqiang, et al.
Veröffentlicht: (2025)
von: Chen, Songqiang, et al.
Veröffentlicht: (2025)
Zero-shot Evaluation of Deep Learning for Java Code Clone Detection
von: Heinze, Thomas S.
Veröffentlicht: (2026)
von: Heinze, Thomas S.
Veröffentlicht: (2026)
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
What Builds Effective In-Context Examples for Code Generation?
von: Li, Dongze, et al.
Veröffentlicht: (2025)
von: Li, Dongze, et al.
Veröffentlicht: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
Compilation of Commit Changes within Java Source Code Repositories
von: Schott, Stefan, et al.
Veröffentlicht: (2024)
von: Schott, Stefan, et al.
Veröffentlicht: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
MapReplay: Trace-Driven Benchmark Generation for Java HashMap
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
JEDI: Java Evaluation of Declarative and Imperative Queries
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
Strengthening Programming Comprehension in Large Language Models through Code Generation
von: Ren, Xiaoning, et al.
Veröffentlicht: (2025)
von: Ren, Xiaoning, et al.
Veröffentlicht: (2025)
Embracing Objects Over Statics: An Analysis of Method Preferences in Open Source Java Frameworks
von: Zakharov, Vladimir, et al.
Veröffentlicht: (2024)
von: Zakharov, Vladimir, et al.
Veröffentlicht: (2024)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
von: Jiang, Mingchao, et al.
Veröffentlicht: (2025)
von: Jiang, Mingchao, et al.
Veröffentlicht: (2025)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
von: Dau, Anh T. V., et al.
Veröffentlicht: (2026)
von: Dau, Anh T. V., et al.
Veröffentlicht: (2026)
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
von: Qiu, Ruidi, et al.
Veröffentlicht: (2024)
von: Qiu, Ruidi, et al.
Veröffentlicht: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
iScript: A Domain-Adapted Large Language Model and Benchmark for Physical Design Tcl Script Generation
von: Xu, Ning, et al.
Veröffentlicht: (2026)
von: Xu, Ning, et al.
Veröffentlicht: (2026)
From Effectiveness to Efficiency: Uncovering Linguistic Bias in Large Language Model-based Code Generation
von: Jiang, Weipeng, et al.
Veröffentlicht: (2024)
von: Jiang, Weipeng, et al.
Veröffentlicht: (2024)
JustinANN: Realistic Test Generation for Java Programs Driven by Annotations
von: Cui, Baoquan, et al.
Veröffentlicht: (2025)
von: Cui, Baoquan, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
von: Wallraven, Stephan, et al.
Veröffentlicht: (2026)
von: Wallraven, Stephan, et al.
Veröffentlicht: (2026)
CodeMind: Evaluating Large Language Models for Code Reasoning
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
von: Liu, Changshu, et al.
Veröffentlicht: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024)
von: Zhang, William, et al.
Veröffentlicht: (2024)
MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing
von: Xu, Congying, et al.
Veröffentlicht: (2024)
von: Xu, Congying, et al.
Veröffentlicht: (2024)
When LLMs Meet API Documentation: Can Retrieval Augmentation Aid Code Generation Just as It Helps Developers?
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
von: Lang, Nguyet-Anh H., et al.
Veröffentlicht: (2026)
Misleading Microbenchmarks on the Java Virtual Machines
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
von: Schiavio, Filippo, et al.
Veröffentlicht: (2026)
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
von: Zhu, Qiming, et al.
Veröffentlicht: (2024)
von: Zhu, Qiming, et al.
Veröffentlicht: (2024)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
von: Fang, Sen, et al.
Veröffentlicht: (2025)
von: Fang, Sen, et al.
Veröffentlicht: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
von: Dandamudi, Rohit, et al.
Veröffentlicht: (2024)
LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators
von: Liu, Yuchen, et al.
Veröffentlicht: (2025)
von: Liu, Yuchen, et al.
Veröffentlicht: (2025)
Scalable, Validated Code Translation of Entire Projects using Large Language Models
von: Zhang, Hanliang, et al.
Veröffentlicht: (2024)
von: Zhang, Hanliang, et al.
Veröffentlicht: (2024)
Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language Benchmarks
von: Nyamsuren, Enkhbold
Veröffentlicht: (2024)
von: Nyamsuren, Enkhbold
Veröffentlicht: (2024)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
Pattern-Based Peephole Optimizations with Java JIT Tests
von: Zang, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Zang, Zhiqiang, et al.
Veröffentlicht: (2024)
Finding Compiler Bugs through Cross-Language Code Generator and Differential Testing
von: Feng, Qiong, et al.
Veröffentlicht: (2025)
von: Feng, Qiong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
von: Wu, Jiarong, et al.
Veröffentlicht: (2025) -
Can Large Language Models Model Programs Formally?
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026) -
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
von: Chen, Zhiyong, et al.
Veröffentlicht: (2026) -
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025) -
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
von: Chen, Songqiang, et al.
Veröffentlicht: (2025)