Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Majdinasab, Vahid, Nikanjam, Amin, Khomh, Foutse |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025)
Fault Localization in Deep Learning-based Software: A System-level Approach
von: Morovati, Mohammad Mehdi, et al.
Veröffentlicht: (2024)
von: Morovati, Mohammad Mehdi, et al.
Veröffentlicht: (2024)
An Efficient Model Maintenance Approach for MLOps
von: Majidi, Forough, et al.
Veröffentlicht: (2024)
von: Majidi, Forough, et al.
Veröffentlicht: (2024)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
ReCatcher: Towards LLMs Regression Testing for Code Generation
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2025)
Bugs in Large Language Models Generated Code: An Empirical Study
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
von: Tambon, Florian, et al.
Veröffentlicht: (2024)
Quality Issues in Machine Learning Software Systems
von: Côté, Pierre-Olivier, et al.
Veröffentlicht: (2023)
von: Côté, Pierre-Olivier, et al.
Veröffentlicht: (2023)
On the Effectiveness of Log Representation for Log-based Anomaly Detection
von: Wu, Xingfang, et al.
Veröffentlicht: (2023)
von: Wu, Xingfang, et al.
Veröffentlicht: (2023)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
von: Bouchoucha, Rached, et al.
Veröffentlicht: (2024)
von: Bouchoucha, Rached, et al.
Veröffentlicht: (2024)
GIST: Generated Inputs Sets Transferability in Deep Learning
von: Tambon, Florian, et al.
Veröffentlicht: (2023)
von: Tambon, Florian, et al.
Veröffentlicht: (2023)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
von: Wu, Xingfang, et al.
Veröffentlicht: (2024)
von: Wu, Xingfang, et al.
Veröffentlicht: (2024)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
Machine Learning Robustness: A Primer
von: Braiek, Houssem Ben, et al.
Veröffentlicht: (2024)
von: Braiek, Houssem Ben, et al.
Veröffentlicht: (2024)
Continuously Learning Bug Locations
von: Mindom, Paulina Stevia Nouwou, et al.
Veröffentlicht: (2024)
von: Mindom, Paulina Stevia Nouwou, et al.
Veröffentlicht: (2024)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
von: Morovati, Mohammad Mehdi, et al.
Veröffentlicht: (2023)
von: Morovati, Mohammad Mehdi, et al.
Veröffentlicht: (2023)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Mining Action Rules for Defect Reduction Planning
von: Oueslati, Khouloud, et al.
Veröffentlicht: (2024)
von: Oueslati, Khouloud, et al.
Veröffentlicht: (2024)
Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study
von: Shah, Mehil B., et al.
Veröffentlicht: (2024)
von: Shah, Mehil B., et al.
Veröffentlicht: (2024)
FairFLRep: Fairness aware fault localization and repair of Deep Neural Networks
von: Openja, Moses, et al.
Veröffentlicht: (2025)
von: Openja, Moses, et al.
Veröffentlicht: (2025)
Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
von: Voria, Gianmario, et al.
Veröffentlicht: (2026)
von: Voria, Gianmario, et al.
Veröffentlicht: (2026)
Trimming the Risk: Towards Reliable Continuous Training for Deep Learning Inspection Systems
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2024)
von: Abbassi, Altaf Allah, et al.
Veröffentlicht: (2024)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
von: Wu, Xingfang, et al.
Veröffentlicht: (2023)
von: Wu, Xingfang, et al.
Veröffentlicht: (2023)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
Calibration and Correctness of Language Models for Code
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
von: Shah, Mehil B, et al.
Veröffentlicht: (2025)
von: Shah, Mehil B, et al.
Veröffentlicht: (2025)
Improving the Robustness of Large Language Models for Code Tasks via Fine-tuning with Perturbed Data
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
von: Taraghi, Mina, et al.
Veröffentlicht: (2024)
von: Taraghi, Mina, et al.
Veröffentlicht: (2024)
Pre-Training Representations of Binary Code Using Contrastive Learning
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
von: Paul, Indraneil, et al.
Veröffentlicht: (2026)
von: Paul, Indraneil, et al.
Veröffentlicht: (2026)
Unlearning Trojans in Large Language Models: A Comparison Between Natural Language and Source Code
von: Kazemi, Mahdi, et al.
Veröffentlicht: (2024)
von: Kazemi, Mahdi, et al.
Veröffentlicht: (2024)
On Trojan Signatures in Large Language Models of Code
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
On Fixing Insecure AI-Generated Code through Model Fine-Tuning and Prompting Strategies
von: Jahromi, Ali Soltanian Fard, et al.
Veröffentlicht: (2026)
von: Jahromi, Ali Soltanian Fard, et al.
Veröffentlicht: (2026)
An empirical study of testing machine learning in the wild
von: Openja, Moses, et al.
Veröffentlicht: (2023)
von: Openja, Moses, et al.
Veröffentlicht: (2023)
Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
An Empirical Study of Policy-as-Code Adoption in Open-Source Software Projects
von: Foalem, Patrick Loic, et al.
Veröffentlicht: (2026)
von: Foalem, Patrick Loic, et al.
Veröffentlicht: (2026)
Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
von: Hussain, Aftab, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024) -
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2025) -
Fault Localization in Deep Learning-based Software: A System-level Approach
von: Morovati, Mohammad Mehdi, et al.
Veröffentlicht: (2024) -
An Efficient Model Maintenance Approach for MLOps
von: Majidi, Forough, et al.
Veröffentlicht: (2024) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
von: Tambon, Florian, et al.
Veröffentlicht: (2024)