Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jahanshahi, Mahmoud, Mockus, Audris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dataset: Copy-based Reuse in Open Source Software
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2023)
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2023)
OSS License Identification at Scale: A Comprehensive Dataset Using World of Code
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024)
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024)
Beyond Dependencies: The Role of Copy-Based Reuse in Open Source Software Development
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024)
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024)
Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software
von: Thakur, Addi Malviya, et al.
Veröffentlicht: (2025)
von: Thakur, Addi Malviya, et al.
Veröffentlicht: (2025)
The Role of Data Filtering in Open Source Software Ranking and Selection
von: Malviya-Thakur, Addi, et al.
Veröffentlicht: (2024)
von: Malviya-Thakur, Addi, et al.
Veröffentlicht: (2024)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
von: Tan, Jingwen, et al.
Veröffentlicht: (2024)
DevLicOps: A Framework for Mitigating Licensing Risks in AI-Generated Code
von: Sharma, Pratyush Nidhi, et al.
Veröffentlicht: (2025)
von: Sharma, Pratyush Nidhi, et al.
Veröffentlicht: (2025)
The Foundation Cracks: A Comprehensive Study on Bugs and Testing Practices in LLM Libraries
von: Jiang, Weipeng, et al.
Veröffentlicht: (2025)
von: Jiang, Weipeng, et al.
Veröffentlicht: (2025)
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair
von: Khan, Zanis Ali, et al.
Veröffentlicht: (2025)
von: Khan, Zanis Ali, et al.
Veröffentlicht: (2025)
Hidden Licensing Risks in the LLMware Ecosystem
von: Wang, Bo, et al.
Veröffentlicht: (2026)
von: Wang, Bo, et al.
Veröffentlicht: (2026)
On the Standardization of Behavioral Use Clauses and Their Adoption for Responsible Licensing of AI
von: McDuff, Daniel, et al.
Veröffentlicht: (2024)
von: McDuff, Daniel, et al.
Veröffentlicht: (2024)
Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection
von: Lbath, Amine
Veröffentlicht: (2026)
von: Lbath, Amine
Veröffentlicht: (2026)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
The Code Whisperer: LLM and Graph-Based AI for Smell and Vulnerability Resolution
von: Baqar, Mohammad, et al.
Veröffentlicht: (2026)
von: Baqar, Mohammad, et al.
Veröffentlicht: (2026)
LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2024)
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2024)
FullStack Bench: Evaluating LLMs as Full Stack Coders
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary Study
von: Antal, Gábor, et al.
Veröffentlicht: (2025)
von: Antal, Gábor, et al.
Veröffentlicht: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
von: Stalnaker, Trevor, et al.
Veröffentlicht: (2024)
von: Stalnaker, Trevor, et al.
Veröffentlicht: (2024)
The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
von: Pan, Dangfeng, et al.
Veröffentlicht: (2025)
von: Pan, Dangfeng, et al.
Veröffentlicht: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
von: Tóth, Rebeka, et al.
Veröffentlicht: (2024)
von: Tóth, Rebeka, et al.
Veröffentlicht: (2024)
Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
von: Du, Xueying, et al.
Veröffentlicht: (2024)
von: Du, Xueying, et al.
Veröffentlicht: (2024)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
von: Jewitt, James, et al.
Veröffentlicht: (2025)
von: Jewitt, James, et al.
Veröffentlicht: (2025)
CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
von: Bifolco, Daniele, et al.
Veröffentlicht: (2025)
von: Bifolco, Daniele, et al.
Veröffentlicht: (2025)
Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics Capacities
von: Ma, Wei, et al.
Veröffentlicht: (2022)
von: Ma, Wei, et al.
Veröffentlicht: (2022)
CrackMeBench: Binary Reverse Engineering for Agents
von: David, Isaac, et al.
Veröffentlicht: (2026)
von: David, Isaac, et al.
Veröffentlicht: (2026)
DRS-OSS: Practical Diff Risk Scoring with LLMs
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2025)
von: Sayedsalehi, Ali, et al.
Veröffentlicht: (2025)
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
von: Kabir, Samia, et al.
Veröffentlicht: (2023)
von: Kabir, Samia, et al.
Veröffentlicht: (2023)
An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process
von: Synovic, Nicholas M., et al.
Veröffentlicht: (2026)
von: Synovic, Nicholas M., et al.
Veröffentlicht: (2026)
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2025)
von: Widyasari, Ratnadira, et al.
Veröffentlicht: (2025)
Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions
von: Wang, Chengjie, et al.
Veröffentlicht: (2026)
von: Wang, Chengjie, et al.
Veröffentlicht: (2026)
An Empirical Evaluation of LLM-Based Approaches for Code Vulnerability Detection: RAG, SFT, and Dual-Agent Systems
von: Saju, Md Hasan, et al.
Veröffentlicht: (2026)
von: Saju, Md Hasan, et al.
Veröffentlicht: (2026)
Cracking CodeWhisperer: Analyzing Developers' Interactions and Patterns During Programming Tasks
von: Javahar, Jeena, et al.
Veröffentlicht: (2025)
von: Javahar, Jeena, et al.
Veröffentlicht: (2025)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
von: Da Silva, Leuson, et al.
Veröffentlicht: (2024)
von: Da Silva, Leuson, et al.
Veröffentlicht: (2024)
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
von: Filho, Elzo Brito dos Santos
Veröffentlicht: (2026)
von: Filho, Elzo Brito dos Santos
Veröffentlicht: (2026)
Software Dependencies 2.0: An Empirical Study of Reuse and Integration of Pre-Trained Models in Open-Source Projects
von: Yasmin, Jerin, et al.
Veröffentlicht: (2025)
von: Yasmin, Jerin, et al.
Veröffentlicht: (2025)
Studying Vulnerable Code Entities in R
von: Zhao, Zixiao, et al.
Veröffentlicht: (2024)
von: Zhao, Zixiao, et al.
Veröffentlicht: (2024)
Moving Faster and Reducing Risk: Using LLMs in Release Deployment
von: Abreu, Rui, et al.
Veröffentlicht: (2024)
von: Abreu, Rui, et al.
Veröffentlicht: (2024)
An Empirical Study of OpenAI API Discussions on Stack Overflow
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
StarCoder 2 and The Stack v2: The Next Generation
von: Lozhkov, Anton, et al.
Veröffentlicht: (2024)
von: Lozhkov, Anton, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dataset: Copy-based Reuse in Open Source Software
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2023) -
OSS License Identification at Scale: A Comprehensive Dataset Using World of Code
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024) -
Beyond Dependencies: The Role of Copy-Based Reuse in Open Source Software Development
von: Jahanshahi, Mahmoud, et al.
Veröffentlicht: (2024) -
Scientific Open-Source Software Is Less Likely to Become Abandoned Than One Might Think! Lessons from Curating a Catalog of Maintained Scientific Software
von: Thakur, Addi Malviya, et al.
Veröffentlicht: (2025) -
The Role of Data Filtering in Open Source Software Ranking and Selection
von: Malviya-Thakur, Addi, et al.
Veröffentlicht: (2024)