Machine Learning in the Wild: Early Evidence of Non-Compliant ML-Automation in Open-Source Software
Fuente:
arXiv
Saved in:
| Main Authors: | Arshid, Zohaib, Bifolco, Daniele, Zampetti, Fiorella, Di Penta, Massimiliano |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot
by: Bifolco, Daniele, et al.
Published: (2025)
by: Bifolco, Daniele, et al.
Published: (2025)
CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
by: Bifolco, Daniele, et al.
Published: (2025)
by: Bifolco, Daniele, et al.
Published: (2025)
How are MLOps Frameworks Used in Open Source Projects? An Empirical Characterization
by: Zampetti, Fiorella, et al.
Published: (2026)
by: Zampetti, Fiorella, et al.
Published: (2026)
Developers and Generative AI: A Study of Self-Admitted Usage in Open Source Projects
by: Tufano, Rosalia, et al.
Published: (2026)
by: Tufano, Rosalia, et al.
Published: (2026)
A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems
by: Pepe, Federica, et al.
Published: (2024)
by: Pepe, Federica, et al.
Published: (2024)
Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization
by: Midolo, Alessandro, et al.
Published: (2026)
by: Midolo, Alessandro, et al.
Published: (2026)
Automated Refactoring of Non-Idiomatic Python Code: A Differentiated Replication with LLMs
by: Midolo, Alessandro, et al.
Published: (2025)
by: Midolo, Alessandro, et al.
Published: (2025)
Augmenting Software Bills of Materials with Software Vulnerability Description: A Preliminary Study on GitHub
by: Fucci, Davide, et al.
Published: (2025)
by: Fucci, Davide, et al.
Published: (2025)
From Human to Machine Refactoring: Assessing GPT-4's Impact on Python Class Quality and Readability
by: Midolo, Alessandro, et al.
Published: (2026)
by: Midolo, Alessandro, et al.
Published: (2026)
Unveiling ChatGPT's Usage in Open Source Projects: A Mining-based Study
by: Tufano, Rosalia, et al.
Published: (2024)
by: Tufano, Rosalia, et al.
Published: (2024)
How the Training Procedure Impacts the Performance of Deep Learning-based Vulnerability Patching
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems
by: Stalnaker, Trevor, et al.
Published: (2023)
by: Stalnaker, Trevor, et al.
Published: (2023)
Future of Software Engineering Research: The SIGSOFT Perspective
by: Di Penta, Massimiliano, et al.
Published: (2026)
by: Di Penta, Massimiliano, et al.
Published: (2026)
An Analysis of Malicious Packages in Open-Source Software in the Wild
by: Zhou, Xiaoyan, et al.
Published: (2024)
by: Zhou, Xiaoyan, et al.
Published: (2024)
"Let it be Chaos in the Plumbing!" Usage and Efficacy of Chaos Engineering in DevOps Pipelines
by: Fossati, Stefano, et al.
Published: (2025)
by: Fossati, Stefano, et al.
Published: (2025)
Automatic Categorization of GitHub Actions with Transformers and Few-shot Learning
by: Nguyen, Phuong T., et al.
Published: (2024)
by: Nguyen, Phuong T., et al.
Published: (2024)
Automated Extraction and Analysis of Developer's Rationale in Open Source Software
by: Dhaouadi, Mouna, et al.
Published: (2025)
by: Dhaouadi, Mouna, et al.
Published: (2025)
Developers' Perspectives on Software Licensing: Current Practices, Challenges, and Tools
by: Wintersgill, Nathan, et al.
Published: (2025)
by: Wintersgill, Nathan, et al.
Published: (2025)
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
by: Giagnorio, Alessandro, et al.
Published: (2025)
by: Giagnorio, Alessandro, et al.
Published: (2025)
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
Self-Admitted Technical Debt in LLM Software: An Empirical Comparison with ML and Non-ML Software
by: Selvanayagam, Niruthiha, et al.
Published: (2026)
by: Selvanayagam, Niruthiha, et al.
Published: (2026)
Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
by: Afrin, Saima, et al.
Published: (2026)
by: Afrin, Saima, et al.
Published: (2026)
An Empirical Analysis of Machine Learning Model and Dataset Documentation, Supply Chain, and Licensing Challenges on Hugging Face
by: Stalnaker, Trevor, et al.
Published: (2025)
by: Stalnaker, Trevor, et al.
Published: (2025)
CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
by: Yu, Zhengmin, et al.
Published: (2025)
by: Yu, Zhengmin, et al.
Published: (2025)
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
by: Stalnaker, Trevor, et al.
Published: (2024)
by: Stalnaker, Trevor, et al.
Published: (2024)
Measuring Software Development Waste in Open-Source Software Projects
by: Varanasi, Dhiraj SM, et al.
Published: (2024)
by: Varanasi, Dhiraj SM, et al.
Published: (2024)
Measuring Software Innovation with Open Source Software Development Data
by: Brown, Eva Maxfield, et al.
Published: (2024)
by: Brown, Eva Maxfield, et al.
Published: (2024)
Unlocking Reproducibility: Automating re-Build Process for Open-Source Software
by: Hassanshahi, Behnaz, et al.
Published: (2025)
by: Hassanshahi, Behnaz, et al.
Published: (2025)
AWML: An Open-Source ML-based Robotics Perception Framework to Deploy for ROS-based Autonomous Driving Software
by: Tanaka, Satoshi, et al.
Published: (2025)
by: Tanaka, Satoshi, et al.
Published: (2025)
Understanding Underrepresented Groups in Open Source Software
by: Santos, Reydne, et al.
Published: (2025)
by: Santos, Reydne, et al.
Published: (2025)
Extension Decisions in Open Source Software Ecosystem
by: Onagh, Elmira, et al.
Published: (2025)
by: Onagh, Elmira, et al.
Published: (2025)
Many-Objective Optimization of Non-Functional Attributes based on Refactoring of Software Models
by: Cortellessa, Vittorio, et al.
Published: (2023)
by: Cortellessa, Vittorio, et al.
Published: (2023)
Interactive GDPR-Compliant Privacy Policy Generation for Software Applications
by: Sangaroonsilp, Pattaraporn, et al.
Published: (2024)
by: Sangaroonsilp, Pattaraporn, et al.
Published: (2024)
Game Elements to Engage Students Learning the Open Source Software Contribution Process
by: Santos, Italo, et al.
Published: (2024)
by: Santos, Italo, et al.
Published: (2024)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
by: González, Alexandra, et al.
Published: (2024)
by: González, Alexandra, et al.
Published: (2024)
Community Engagement and the Lifespan of Open-Source Software Projects
by: Kaushik, Mohit, et al.
Published: (2025)
by: Kaushik, Mohit, et al.
Published: (2025)
Context Engineering for AI Agents in Open-Source Software
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
by: Mohsenimofidi, Seyedmoein, et al.
Published: (2025)
Dataset: Copy-based Reuse in Open Source Software
by: Jahanshahi, Mahmoud, et al.
Published: (2023)
by: Jahanshahi, Mahmoud, et al.
Published: (2023)
The Product Beyond the Model -- An Empirical Study of Repositories of Open-Source ML Products
by: Nahar, Nadia, et al.
Published: (2023)
by: Nahar, Nadia, et al.
Published: (2023)
Don't Disturb Me: Challenges of Interacting with SoftwareBots on Open Source Software Projects
by: Wessel, Mairieli, et al.
Published: (2021)
by: Wessel, Mairieli, et al.
Published: (2021)
Similar Items
-
Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot
by: Bifolco, Daniele, et al.
Published: (2025) -
CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
by: Bifolco, Daniele, et al.
Published: (2025) -
How are MLOps Frameworks Used in Open Source Projects? An Empirical Characterization
by: Zampetti, Fiorella, et al.
Published: (2026) -
Developers and Generative AI: A Study of Self-Admitted Usage in Open Source Projects
by: Tufano, Rosalia, et al.
Published: (2026) -
A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems
by: Pepe, Federica, et al.
Published: (2024)