On Inter-dataset Code Duplication and Data Leakage in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | López, José Antonio Hernández, Chen, Boqi, Saaz, Mootez, Sharma, Tushar, Varró, Dániel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
ALPINE: An adaptive language-agnostic pruning method for language models for code
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
CONCORD: Towards a DSL for Configurable Graph Code Representation
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
SHERPA: A Model-Driven Framework for Large Language Model Execution
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
On the Effect of Token Merging on Pre-trained Models for Code
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
von: Saad, Mootez, et al.
Veröffentlicht: (2025)
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
Tu(r)ning AI Green: Exploring Energy Efficiency Cascading with Orthogonal Optimizations
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2025)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2025)
LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker Generation
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
Generative AI in Simulation-Based Test Environments for Large-Scale Cyber-Physical Systems: An Industrial Study
von: Sadrnezhaad, Masoud, et al.
Veröffentlicht: (2025)
von: Sadrnezhaad, Masoud, et al.
Veröffentlicht: (2025)
CodeMorph: Mitigating Data Leakage in Large Language Model Assessment
von: Rao, Hongzhou, et al.
Veröffentlicht: (2025)
von: Rao, Hongzhou, et al.
Veröffentlicht: (2025)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
von: Wang, Yiran, et al.
Veröffentlicht: (2025)
Runtime-Augmented LLMs for Crash Detection and Diagnosis in ML Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2026)
von: Wang, Yiran, et al.
Veröffentlicht: (2026)
Why do Machine Learning Notebooks Crash? An Empirical Study on Public Python Jupyter Notebooks
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
MCeT: Behavioral Model Correctness Evaluation using Large Language Models
von: Ahmed, Khaled, et al.
Veröffentlicht: (2025)
von: Ahmed, Khaled, et al.
Veröffentlicht: (2025)
CodeGreen: Towards Improving Precision and Portability in Software Energy Measurement
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
Projectional Decoding: Towards Semantic-Aware LLM Generation
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
Structure- and Event-Driven Frameworks for State Machine Modeling with Large Language Models
von: Abdulkarim, Samer, et al.
Veröffentlicht: (2026)
von: Abdulkarim, Samer, et al.
Veröffentlicht: (2026)
Accurate and Consistent Graph Model Generation from Text with Large Language Models
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
von: Chen, Boqi, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
von: Afrin, Saima, et al.
Veröffentlicht: (2026)
von: Afrin, Saima, et al.
Veröffentlicht: (2026)
Energy Flow Graph: Modeling Software Energy Consumption
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
von: Rajput, Saurabhsingh, et al.
Veröffentlicht: (2026)
Concretization of Abstract Traffic Scene Specifications Using Metaheuristic Search
von: Babikian, Aren A., et al.
Veröffentlicht: (2023)
von: Babikian, Aren A., et al.
Veröffentlicht: (2023)
Broken Windows: Exploring the Applicability of a Controversial Theory on Code Quality
von: Spinellis, Diomidis, et al.
Veröffentlicht: (2024)
von: Spinellis, Diomidis, et al.
Veröffentlicht: (2024)
Generating refactored code accurately using reinforcement learning
von: Palit, Indranil, et al.
Veröffentlicht: (2024)
von: Palit, Indranil, et al.
Veröffentlicht: (2024)
An Empirical Study on the Impact of Code Duplication-aware Refactoring Practices on Quality Metrics
von: AlOmar, Eman Abdullah
Veröffentlicht: (2025)
von: AlOmar, Eman Abdullah
Veröffentlicht: (2025)
Aligning Requirement for Large Language Model's Code Generation
von: Tian, Zhao, et al.
Veröffentlicht: (2025)
von: Tian, Zhao, et al.
Veröffentlicht: (2025)
Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2025)
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2025)
Enhanced Automated Code Vulnerability Repair using Large Language Models
von: de-Fitero-Dominguez, David, et al.
Veröffentlicht: (2024)
von: de-Fitero-Dominguez, David, et al.
Veröffentlicht: (2024)
The Impact of Large Language Models (LLMs) on Code Review Process
von: Collante, Antonio, et al.
Veröffentlicht: (2025)
von: Collante, Antonio, et al.
Veröffentlicht: (2025)
Bogus Bugs, Duplicates, and Revealing Comments: Data Quality Issues in NPR
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
Hotfixing Large Language Models for Code
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
Ecosystem of Large Language Models for Code
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
von: Yang, Zhou, et al.
Veröffentlicht: (2024)
Greening Large Language Models of Code
von: Shi, Jieke, et al.
Veröffentlicht: (2023)
von: Shi, Jieke, et al.
Veröffentlicht: (2023)
Rethinking Code Complexity Through the Lens of Large Language Models
von: Xie, Chen, et al.
Veröffentlicht: (2026)
von: Xie, Chen, et al.
Veröffentlicht: (2026)
Advancing Code Coverage: Incorporating Program Analysis with Large Language Models
von: Yang, Chen, et al.
Veröffentlicht: (2024)
von: Yang, Chen, et al.
Veröffentlicht: (2024)
LeakageDetector: An Open Source Data Leakage Analysis Tool in Machine Learning Pipelines
von: AlOmar, Eman Abdullah, et al.
Veröffentlicht: (2025)
von: AlOmar, Eman Abdullah, et al.
Veröffentlicht: (2025)
DCE-LLM: Dead Code Elimination with Large Language Models
von: Chen, Minyu, et al.
Veröffentlicht: (2025)
von: Chen, Minyu, et al.
Veröffentlicht: (2025)
An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
von: Xing, Chengli, et al.
Veröffentlicht: (2026)
von: Xing, Chengli, et al.
Veröffentlicht: (2026)
A Validated Taxonomy on Software Energy Smells
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2026)
von: Mehditabar, Mohammadjavad, et al.
Veröffentlicht: (2026)
Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
von: Tian, Zhao, et al.
Veröffentlicht: (2023)
von: Tian, Zhao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
von: Saad, Mootez, et al.
Veröffentlicht: (2025) -
ALPINE: An adaptive language-agnostic pruning method for language models for code
von: Saad, Mootez, et al.
Veröffentlicht: (2024) -
SENAI: Towards Software Engineering Native Generative Artificial Intelligence
von: Saad, Mootez, et al.
Veröffentlicht: (2025) -
CONCORD: Towards a DSL for Configurable Graph Code Representation
von: Saad, Mootez, et al.
Veröffentlicht: (2024) -
SHERPA: A Model-Driven Framework for Large Language Model Execution
von: Chen, Boqi, et al.
Veröffentlicht: (2025)