Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Saad, Mootez, Chen, Boqi, López, José Antonio Hernández, Varró, Dániel, Sharma, Tushar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Inter-dataset Code Duplication and Data Leakage in Large Language Models
von: López, José Antonio Hernández, et al.
Veröffentlicht: (2024)
von: López, José Antonio Hernández, et al.
Veröffentlicht: (2024)
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
ALPINE: An adaptive language-agnostic pruning method for language models for code
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
CONCORD: Towards a DSL for Configurable Graph Code Representation
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
On the Effect of Token Merging on Pre-trained Models for Code
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
Tu(r)ning AI Green: Exploring Energy Efficiency Cascading with Orthogonal Optimizations
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2025)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2025)
SHERPA: A Model-Driven Framework for Large Language Model Execution
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker Generation
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
Generative AI in Simulation-Based Test Environments for Large-Scale Cyber-Physical Systems: An Industrial Study
von: Sadrnezhaad, Masoud, et al.
Veröffentlicht: (2025)
von: Sadrnezhaad, Masoud, et al.
Veröffentlicht: (2025)
CodeGreen: Towards Improving Precision and Portability in Software Energy Measurement
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
MCeT: Behavioral Model Correctness Evaluation using Large Language Models
von: Ahmed, Khaled, et al.
Veröffentlicht: (2025)
von: Ahmed, Khaled, et al.
Veröffentlicht: (2025)
Energy Flow Graph: Modeling Software Energy Consumption
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
Runtime-Augmented LLMs for Crash Detection and Diagnosis in ML Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2026)
von: Wang, Yiran, et al.
Veröffentlicht: (2026)
Why do Machine Learning Notebooks Crash? An Empirical Study on Public Python Jupyter Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
A Validated Taxonomy on Software Energy Smells
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2026)
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2026)
Projectional Decoding: Towards Semantic-Aware LLM Generation
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
Structure- and Event-Driven Frameworks for State Machine Modeling with Large Language Models
von: Abdulkarim, Samer, et al.
Veröffentlicht: (2026)
von: Abdulkarim, Samer, et al.
Veröffentlicht: (2026)
Accurate and Consistent Graph Model Generation from Text with Large Language Models
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
An Empirical Study on the Code Refactoring Capability of Large Language Models
von: Cordeiro, Jonathan, et al.
Veröffentlicht: (2024)
von: Cordeiro, Jonathan, et al.
Veröffentlicht: (2024)
Measuring Determinism in Large Language Models for Software Code Review
von: Klishevich, Eugene, et al.
Veröffentlicht: (2025)
von: Klishevich, Eugene, et al.
Veröffentlicht: (2025)
Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
von: He, Mengliang, et al.
Veröffentlicht: (2025)
von: He, Mengliang, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Code Translation: Effects of Prompt Language and Prompt Design
von: Aljagthami, Aamer, et al.
Veröffentlicht: (2025)
von: Aljagthami, Aamer, et al.
Veröffentlicht: (2025)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
von: Li, Linyi, et al.
Veröffentlicht: (2024)
von: Li, Linyi, et al.
Veröffentlicht: (2024)
A Pilot Study on Detecting Software Design Patterns with Large Language Models: An Empirical Evaluation
von: Chowdhury, Oishik, et al.
Veröffentlicht: (2026)
von: Chowdhury, Oishik, et al.
Veröffentlicht: (2026)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
von: Yang, Rem, et al.
Veröffentlicht: (2025)
von: Yang, Rem, et al.
Veröffentlicht: (2025)
Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
von: Velasco, Alejandro, et al.
Veröffentlicht: (2024)
von: Velasco, Alejandro, et al.
Veröffentlicht: (2024)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
von: Padwal, Vedant
Veröffentlicht: (2026)
von: Padwal, Vedant
Veröffentlicht: (2026)
Concretization of Abstract Traffic Scene Specifications Using Metaheuristic Search
von: Babikian, Aren A., et al.
Veröffentlicht: (2023)
von: Babikian, Aren A., et al.
Veröffentlicht: (2023)
Broken Windows: Exploring the Applicability of a Controversial Theory on Code Quality
von: Spinellis, Diomidis, et al.
Veröffentlicht: (2024)
von: Spinellis, Diomidis, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
von: Giagnorio, Alessandro, et al.
Veröffentlicht: (2025)
von: Giagnorio, Alessandro, et al.
Veröffentlicht: (2025)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
von: Hu, Xing, et al.
Veröffentlicht: (2025)
von: Hu, Xing, et al.
Veröffentlicht: (2025)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
Assessing the Code Clone Detection Capability of Large Language Models
von: Zhang, Zixian, et al.
Veröffentlicht: (2024)
von: Zhang, Zixian, et al.
Veröffentlicht: (2024)
TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models
von: Chen, Yongxi, et al.
Veröffentlicht: (2025)
von: Chen, Yongxi, et al.
Veröffentlicht: (2025)
Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
von: Sun, Zhihong, et al.
Veröffentlicht: (2025)
von: Sun, Zhihong, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
von: Afrin, Saima, et al.
Veröffentlicht: (2026)
von: Afrin, Saima, et al.
Veröffentlicht: (2026)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On Inter-dataset Code Duplication and Data Leakage in Large Language Models
von: López, José Antonio Hernández, et al.
Veröffentlicht: (2024) -
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
von: Saad, Mootez, et al.
Veröffentlicht: (2025) -
ALPINE: An adaptive language-agnostic pruning method for language models for code
von: Saad, Mootez, et al.
Veröffentlicht: (2024) -
CONCORD: Towards a DSL for Configurable Graph Code Representation
von: Saad, Mootez, et al.
Veröffentlicht: (2024) -
On the Effect of Token Merging on Pre-trained Models for Code
von: Saad, Mootez, et al.
Veröffentlicht: (2025)