Studying LLM Performance on Closed- and Open-source Data
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Toufique, Bird, Christian, Devanbu, Premkumar, Chakraborty, Saikat |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
by: Pai, Kunal, et al.
Published: (2025)
by: Pai, Kunal, et al.
Published: (2025)
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
by: Ahmed, Toufique, et al.
Published: (2023)
by: Ahmed, Toufique, et al.
Published: (2023)
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024)
by: Virk, Yuvraj, et al.
Published: (2024)
Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
by: Hussain, Aftab, et al.
Published: (2024)
by: Hussain, Aftab, et al.
Published: (2024)
Towards Understanding What Code Language Models Learned
by: Ahmed, Toufique, et al.
Published: (2023)
by: Ahmed, Toufique, et al.
Published: (2023)
Calibration and Correctness of Language Models for Code
by: Spiess, Claudio, et al.
Published: (2024)
by: Spiess, Claudio, et al.
Published: (2024)
Investigating Test Overfitting on SWE-bench
by: Ahmed, Toufique, et al.
Published: (2025)
by: Ahmed, Toufique, et al.
Published: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
Otter: Generating Tests from Issues to Validate SWE Patches
by: Ahmed, Toufique, et al.
Published: (2025)
by: Ahmed, Toufique, et al.
Published: (2025)
How Robustly do LLMs Understand Execution Semantics?
by: Spiess, Claudio, et al.
Published: (2026)
by: Spiess, Claudio, et al.
Published: (2026)
Ecosystem of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
On LLMs' Internal Representation of Code Correctness
by: Ribeiro, Francisco, et al.
Published: (2025)
by: Ribeiro, Francisco, et al.
Published: (2025)
CoCoNUT: Structural Code Understanding does not fall out of a tree
by: Beger, Claas, et al.
Published: (2025)
by: Beger, Claas, et al.
Published: (2025)
LLM Performance for Code Generation on Noisy Tasks
by: Sendyka, Radzim, et al.
Published: (2025)
by: Sendyka, Radzim, et al.
Published: (2025)
Beyond the Comfort Zone: Emerging Solutions to Overcome Challenges in Integrating LLMs into Software Products
by: Nahar, Nadia, et al.
Published: (2024)
by: Nahar, Nadia, et al.
Published: (2024)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
by: Haider, Md. Asif, et al.
Published: (2024)
by: Haider, Md. Asif, et al.
Published: (2024)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Machine Learning Operations: A Mapping Study
by: Chakraborty, Abhijit, et al.
Published: (2024)
by: Chakraborty, Abhijit, et al.
Published: (2024)
A Survey of Trojans in Neural Models of Source Code: Taxonomy and Techniques
by: Hussain, Aftab, et al.
Published: (2023)
by: Hussain, Aftab, et al.
Published: (2023)
Data Virtualization for Machine Learning
by: Khan, Saiful, et al.
Published: (2025)
by: Khan, Saiful, et al.
Published: (2025)
LLM-Based Design Pattern Detection
by: Schindler, Christian, et al.
Published: (2025)
by: Schindler, Christian, et al.
Published: (2025)
Learning to Parallelize with OpenMP by Augmented Heterogeneous AST Representation
by: Chen, Le, et al.
Published: (2023)
by: Chen, Le, et al.
Published: (2023)
Choose Your Simulator Wisely: A Review on Open-source Simulators for Autonomous Driving
by: Li, Yueyuan, et al.
Published: (2023)
by: Li, Yueyuan, et al.
Published: (2023)
Ontology- and LLM-based Data Harmonization for Federated Learning in Healthcare
by: Kokash, Natallia, et al.
Published: (2025)
by: Kokash, Natallia, et al.
Published: (2025)
Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in Deployment
by: Ahmed, Shibbir, et al.
Published: (2024)
by: Ahmed, Shibbir, et al.
Published: (2024)
RocketPPA: Code-Level Power, Performance, and Area Prediction via LLM and Mixture of Experts
by: Abdollahi, Armin, et al.
Published: (2025)
by: Abdollahi, Armin, et al.
Published: (2025)
Towards Refining Developer Questions using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
by: Fathollahzadeh, Pouya, et al.
Published: (2025)
by: Fathollahzadeh, Pouya, et al.
Published: (2025)
When Less is More: On the Value of "Co-training" for Semi-Supervised Software Defect Predictors
by: Majumder, Suvodeep, et al.
Published: (2022)
by: Majumder, Suvodeep, et al.
Published: (2022)
Towards Causal Deep Learning for Vulnerability Detection
by: Rahman, Md Mahbubur, et al.
Published: (2023)
by: Rahman, Md Mahbubur, et al.
Published: (2023)
pyAKI -- An Open Source Solution to Automated KDIGO classification
by: Porschen, Christian, et al.
Published: (2024)
by: Porschen, Christian, et al.
Published: (2024)
An ML-based Approach to Predicting Software Change Dependencies: Insights from an Empirical Study on OpenStack
by: Arabat, Ali, et al.
Published: (2025)
by: Arabat, Ali, et al.
Published: (2025)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Code Roulette: How Prompt Variability Affects LLM Code Generation
by: Paleyes, Andrei, et al.
Published: (2025)
by: Paleyes, Andrei, et al.
Published: (2025)
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
by: Tsai, Yun-Da, et al.
Published: (2024)
by: Tsai, Yun-Da, et al.
Published: (2024)
Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries Across Software Package Ecosystems
by: Li, Hao, et al.
Published: (2022)
by: Li, Hao, et al.
Published: (2022)
On STPA for Distributed Development of Safe Autonomous Driving: An Interview Study
by: Nouri, Ali, et al.
Published: (2024)
by: Nouri, Ali, et al.
Published: (2024)
LLM Critics Help Catch LLM Bugs
by: McAleese, Nat, et al.
Published: (2024)
by: McAleese, Nat, et al.
Published: (2024)
Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time
by: Popescu, Razvan Mihai, et al.
Published: (2026)
by: Popescu, Razvan Mihai, et al.
Published: (2026)
Similar Items
-
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
by: Pai, Kunal, et al.
Published: (2025) -
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
by: Ahmed, Toufique, et al.
Published: (2023) -
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?
by: Ahmed, Toufique, et al.
Published: (2024) -
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024) -
Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
by: Hussain, Aftab, et al.
Published: (2024)