SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
Fuente:
arXiv
Salvato in:
| Autori principali: | Rodriguez-Cardenas, Daniel, Velasco, Alejandro, Poshyvanyk, Denys |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
di: Palacio, David N., et al.
Pubblicazione: (2024)
di: Palacio, David N., et al.
Pubblicazione: (2024)
Toward a Theory of Causation for Interpreting Neural Code Models
di: Palacio, David N., et al.
Pubblicazione: (2023)
di: Palacio, David N., et al.
Pubblicazione: (2023)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
di: Khati, Dipin, et al.
Pubblicazione: (2025)
di: Khati, Dipin, et al.
Pubblicazione: (2025)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
di: Khati, Dipin, et al.
Pubblicazione: (2026)
di: Khati, Dipin, et al.
Pubblicazione: (2026)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
di: Granger, Cole, et al.
Pubblicazione: (2026)
di: Granger, Cole, et al.
Pubblicazione: (2026)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
di: Yang, Hua, et al.
Pubblicazione: (2025)
di: Yang, Hua, et al.
Pubblicazione: (2025)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
di: Yang, Hua, et al.
Pubblicazione: (2025)
di: Yang, Hua, et al.
Pubblicazione: (2025)
Toward Explaining Large Language Models in Software Engineering Tasks
di: Vitale, Antonio, et al.
Pubblicazione: (2025)
di: Vitale, Antonio, et al.
Pubblicazione: (2025)
On Interpreting the Effectiveness of Unsupervised Software Traceability with Information Theory
di: Palacio, David N., et al.
Pubblicazione: (2024)
di: Palacio, David N., et al.
Pubblicazione: (2024)
Toward Neurosymbolic Program Comprehension
di: Velasco, Alejandro, et al.
Pubblicazione: (2025)
di: Velasco, Alejandro, et al.
Pubblicazione: (2025)
Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
di: Velasco, Alejandro, et al.
Pubblicazione: (2024)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2026)
di: Rodriguez-Cardenas, Daniel, et al.
Pubblicazione: (2026)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
di: Kovrigin, Alexander, et al.
Pubblicazione: (2024)
di: Kovrigin, Alexander, et al.
Pubblicazione: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
di: Alor, Ebube, et al.
Pubblicazione: (2025)
di: Alor, Ebube, et al.
Pubblicazione: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
di: Rahman, Imranur, et al.
Pubblicazione: (2025)
di: Rahman, Imranur, et al.
Pubblicazione: (2025)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
di: Khati, Dipin, et al.
Pubblicazione: (2025)
di: Khati, Dipin, et al.
Pubblicazione: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
di: Vulićević, Jelena Ilić
Pubblicazione: (2026)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
di: Xu, Weiwei, et al.
Pubblicazione: (2024)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
di: Wu, Yixi, et al.
Pubblicazione: (2024)
di: Wu, Yixi, et al.
Pubblicazione: (2024)
On LLMs' Internal Representation of Code Correctness
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
di: Ribeiro, Francisco, et al.
Pubblicazione: (2025)
Operational Robustness of LLMs on Code Generation
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
di: Paul, Debalina Ghosh, et al.
Pubblicazione: (2026)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
di: O'Brien, Conor, et al.
Pubblicazione: (2024)
di: O'Brien, Conor, et al.
Pubblicazione: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
di: Thillen, Alex, et al.
Pubblicazione: (2026)
di: Thillen, Alex, et al.
Pubblicazione: (2026)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
di: Patel, Harsh, et al.
Pubblicazione: (2024)
di: Patel, Harsh, et al.
Pubblicazione: (2024)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
di: Galimzyanov, Timur, et al.
Pubblicazione: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
di: Chakrabarty, Sayak, et al.
Pubblicazione: (2024)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
di: Wang, Juntong, et al.
Pubblicazione: (2026)
di: Wang, Juntong, et al.
Pubblicazione: (2026)
LLMs in Coding and their Impact on the Commercial Software Engineering Landscape
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
di: Belozerov, Vladislav, et al.
Pubblicazione: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
di: Bodla, Krishna Vamshi, et al.
Pubblicazione: (2025)
The Struggles of LLMs in Cross-lingual Code Clone Detection
di: Moumoula, Micheline Bénédicte, et al.
Pubblicazione: (2024)
di: Moumoula, Micheline Bénédicte, et al.
Pubblicazione: (2024)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
di: Balla, Stefano, et al.
Pubblicazione: (2026)
di: Balla, Stefano, et al.
Pubblicazione: (2026)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
di: Hu, Ruida, et al.
Pubblicazione: (2025)
di: Hu, Ruida, et al.
Pubblicazione: (2025)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
di: Abedu, Samuel, et al.
Pubblicazione: (2024)
di: Abedu, Samuel, et al.
Pubblicazione: (2024)
LLM-based Content Classification Approach for GitHub Repositories by the README Files
di: Mehmood, Malik Uzair, et al.
Pubblicazione: (2025)
di: Mehmood, Malik Uzair, et al.
Pubblicazione: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
di: Xiao, Yijia, et al.
Pubblicazione: (2025)
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
di: Velasco, Alejandro, et al.
Pubblicazione: (2025)
di: Velasco, Alejandro, et al.
Pubblicazione: (2025)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
di: Otten, Daniel, et al.
Pubblicazione: (2026)
di: Otten, Daniel, et al.
Pubblicazione: (2026)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
di: Gong, Linyuan, et al.
Pubblicazione: (2024)
Lessons Learned: A Multi-Agent Framework for Code LLMs to Learn and Improve
di: Liu, Yuanzhe, et al.
Pubblicazione: (2025)
di: Liu, Yuanzhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
di: Palacio, David N., et al.
Pubblicazione: (2024) -
Toward a Theory of Causation for Interpreting Neural Code Models
di: Palacio, David N., et al.
Pubblicazione: (2023) -
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
di: Khati, Dipin, et al.
Pubblicazione: (2025) -
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
di: Velasco, Alejandro, et al.
Pubblicazione: (2024) -
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
di: Khati, Dipin, et al.
Pubblicazione: (2026)