Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
Fuente:
arXiv
Salvato in:
| Autori principali: | Mundhra, Yash, Valk, Max, Izadi, Maliheh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Traces of Memorisation in Large Language Models for Code
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023)
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
di: Deljouyi, Amirhossein, et al.
Pubblicazione: (2024)
di: Deljouyi, Amirhossein, et al.
Pubblicazione: (2024)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025)
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
AST-PAC: AST-guided Membership Inference for Code
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
Does In-IDE Calibration of Large Language Models work at Scale?
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
Code4MeV2: a Research-oriented Code-completion Platform
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
di: Cipollone, Daniele, et al.
Pubblicazione: (2025)
Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
di: Koohestani, Roham, et al.
Pubblicazione: (2025)
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
di: Katzy, Jonathan, et al.
Pubblicazione: (2025)
TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
di: Koohestani, Roham, et al.
Pubblicazione: (2026)
Prompt-with-Me: in-IDE Structured Prompt Management for LLM-Driven Software Engineering
di: Li, Ziyou, et al.
Pubblicazione: (2025)
di: Li, Ziyou, et al.
Pubblicazione: (2025)
A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
di: de Moor, Aral, et al.
Pubblicazione: (2024)
di: de Moor, Aral, et al.
Pubblicazione: (2024)
Code Readability in the Age of Large Language Models: An Industrial Case Study from Atlassian
di: Takerngsaksiri, Wannita, et al.
Pubblicazione: (2025)
di: Takerngsaksiri, Wannita, et al.
Pubblicazione: (2025)
Long Code Arena: a Set of Benchmarks for Long-Context Code Models
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
di: Bogomolov, Egor, et al.
Pubblicazione: (2024)
Human-AI Experience in Integrated Development Environments: A Systematic Literature Review
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2025)
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2025)
A Note on Code Quality Score: LLMs for Maintainable Large Codebases
di: Wong, Sherman, et al.
Pubblicazione: (2025)
di: Wong, Sherman, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Code Review
di: Cihan, Umut, et al.
Pubblicazione: (2025)
di: Cihan, Umut, et al.
Pubblicazione: (2025)
Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)
A Multi-agent Onboarding Assistant based on Large Language Models, Retrieval Augmented Generation, and Chain-of-Thought
di: Ionescu, Andrei Cristian, et al.
Pubblicazione: (2025)
di: Ionescu, Andrei Cristian, et al.
Pubblicazione: (2025)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study
di: Chand, Sivajeet, et al.
Pubblicazione: (2026)
di: Chand, Sivajeet, et al.
Pubblicazione: (2026)
How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
di: Padwal, Vedant
Pubblicazione: (2026)
di: Padwal, Vedant
Pubblicazione: (2026)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
di: Popescu, Razvan Mihai, et al.
Pubblicazione: (2026)
di: Popescu, Razvan Mihai, et al.
Pubblicazione: (2026)
Automating Patch Set Generation from Code Review Comments Using Large Language Models
di: Rahman, Tajmilur, et al.
Pubblicazione: (2024)
di: Rahman, Tajmilur, et al.
Pubblicazione: (2024)
Bugs in Large Language Models Generated Code: An Empirical Study
di: Tambon, Florian, et al.
Pubblicazione: (2024)
di: Tambon, Florian, et al.
Pubblicazione: (2024)
Output Format Biases in the Evaluation of Large Language Models for Code Translation
di: Macedo, Marcos, et al.
Pubblicazione: (2024)
di: Macedo, Marcos, et al.
Pubblicazione: (2024)
Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
di: He, Mengliang, et al.
Pubblicazione: (2025)
di: He, Mengliang, et al.
Pubblicazione: (2025)
Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
di: Wang, Yanlin, et al.
Pubblicazione: (2024)
di: Wang, Yanlin, et al.
Pubblicazione: (2024)
CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation
di: Xiong, Zhanhang, et al.
Pubblicazione: (2025)
di: Xiong, Zhanhang, et al.
Pubblicazione: (2025)
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models
di: Meaden, James, et al.
Pubblicazione: (2025)
di: Meaden, James, et al.
Pubblicazione: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
di: Zhang, Tanghaoran, et al.
Pubblicazione: (2026)
di: Zhang, Tanghaoran, et al.
Pubblicazione: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
di: Katzy, Jonathan, et al.
Pubblicazione: (2024)
di: Katzy, Jonathan, et al.
Pubblicazione: (2024)
An Empirical Study on Capability of Large Language Models in Understanding Code Semantics
di: Nguyen, Thu-Trang, et al.
Pubblicazione: (2024)
di: Nguyen, Thu-Trang, et al.
Pubblicazione: (2024)
In-IDE Human-AI Experience in the Era of Large Language Models; A Literature Review
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2024)
di: Sergeyuk, Agnia, et al.
Pubblicazione: (2024)
Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
di: Li, Kefan, et al.
Pubblicazione: (2024)
di: Li, Kefan, et al.
Pubblicazione: (2024)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
Beyond Autoregression: An Empirical Study of Diffusion Large Language Models for Code Generation
di: Li, Chengze, et al.
Pubblicazione: (2025)
di: Li, Chengze, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Traces of Memorisation in Large Language Models for Code
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2023) -
Leveraging Large Language Models for Enhancing the Understandability of Generated Unit Tests
di: Deljouyi, Amirhossein, et al.
Pubblicazione: (2024) -
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2025) -
Rethinking IDE Customization for Enhanced HAX: A Hyperdimensional Perspective
di: Koohestani, Roham, et al.
Pubblicazione: (2025) -
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
di: Al-Kaswan, Ali, et al.
Pubblicazione: (2026)