An Empirical Analysis of Machine Learning Model and Dataset Documentation, Supply Chain, and Licensing Challenges on Hugging Face
Fuente:
arXiv
Guardado en:
| Autores principales: | Stalnaker, Trevor, Wintersgill, Nathan, Chaparro, Oscar, Heymann, Laura A., Di Penta, Massimiliano, German, Daniel M, Poshyvanyk, Denys |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
"The Law Doesn't Work Like a Computer": Exploring Software Licensing Issues Faced by Legal Practitioners
por: Wintersgill, Nathan, et al.
Publicado: (2024)
por: Wintersgill, Nathan, et al.
Publicado: (2024)
Developers' Perspectives on Software Licensing: Current Practices, Challenges, and Tools
por: Wintersgill, Nathan, et al.
Publicado: (2025)
por: Wintersgill, Nathan, et al.
Publicado: (2025)
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
por: Stalnaker, Trevor, et al.
Publicado: (2024)
por: Stalnaker, Trevor, et al.
Publicado: (2024)
BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems
por: Stalnaker, Trevor, et al.
Publicado: (2023)
por: Stalnaker, Trevor, et al.
Publicado: (2023)
Prompting in Practice: Investigating Software Practitioners' Use of Generative AI Tools
por: Otten, Daniel, et al.
Publicado: (2025)
por: Otten, Daniel, et al.
Publicado: (2025)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
por: Otten, Daniel, et al.
Publicado: (2026)
por: Otten, Daniel, et al.
Publicado: (2026)
Challenges and Practices in Quantum Software Testing and Debugging: Insights from Practitioners
por: Zappin, Jake, et al.
Publicado: (2025)
por: Zappin, Jake, et al.
Publicado: (2025)
When Quantum Meets Classical: Characterizing Hybrid Quantum-Classical Issues Discussed in Developer Forums
por: Zappin, Jake, et al.
Publicado: (2024)
por: Zappin, Jake, et al.
Publicado: (2024)
Bridging the Quantum Divide: Aligning Academic and Industry Goals in Software Engineering
por: Zappin, Jake, et al.
Publicado: (2025)
por: Zappin, Jake, et al.
Publicado: (2025)
Semantic GUI Scene Learning and Video Alignment for Detecting Duplicate Video-based Bug Reports
por: Yan, Yanfu, et al.
Publicado: (2024)
por: Yan, Yanfu, et al.
Publicado: (2024)
CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
por: Bifolco, Daniele, et al.
Publicado: (2025)
por: Bifolco, Daniele, et al.
Publicado: (2025)
From Human to Machine Refactoring: Assessing GPT-4's Impact on Python Class Quality and Readability
por: Midolo, Alessandro, et al.
Publicado: (2026)
por: Midolo, Alessandro, et al.
Publicado: (2026)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
por: Mastropaolo, Antonio, et al.
Publicado: (2025)
por: Mastropaolo, Antonio, et al.
Publicado: (2025)
Automated Refactoring of Non-Idiomatic Python Code: A Differentiated Replication with LLMs
por: Midolo, Alessandro, et al.
Publicado: (2025)
por: Midolo, Alessandro, et al.
Publicado: (2025)
On the Generalizability of Transformer Models to Code Completions of Different Lengths
por: Cooper, Nathan, et al.
Publicado: (2025)
por: Cooper, Nathan, et al.
Publicado: (2025)
From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem
por: Jewitt, James, et al.
Publicado: (2025)
por: Jewitt, James, et al.
Publicado: (2025)
An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
por: Liu, Yujian, et al.
Publicado: (2026)
por: Liu, Yujian, et al.
Publicado: (2026)
An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
por: Jia, Nan, et al.
Publicado: (2025)
por: Jia, Nan, et al.
Publicado: (2025)
Machine Learning in the Wild: Early Evidence of Non-Compliant ML-Automation in Open-Source Software
por: Arshid, Zohaib, et al.
Publicado: (2026)
por: Arshid, Zohaib, et al.
Publicado: (2026)
How are MLOps Frameworks Used in Open Source Projects? An Empirical Characterization
por: Zampetti, Fiorella, et al.
Publicado: (2026)
por: Zampetti, Fiorella, et al.
Publicado: (2026)
Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?
por: Vitale, Antonio, et al.
Publicado: (2025)
por: Vitale, Antonio, et al.
Publicado: (2025)
A Large-Scale Exploit Instrumentation Study of AI/ML Supply Chain Attacks in Hugging Face Models
por: Casey, Beatrice, et al.
Publicado: (2024)
por: Casey, Beatrice, et al.
Publicado: (2024)
Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization
por: Midolo, Alessandro, et al.
Publicado: (2026)
por: Midolo, Alessandro, et al.
Publicado: (2026)
Augmenting Software Bills of Materials with Software Vulnerability Description: A Preliminary Study on GitHub
por: Fucci, Davide, et al.
Publicado: (2025)
por: Fucci, Davide, et al.
Publicado: (2025)
How the Training Procedure Impacts the Performance of Deep Learning-based Vulnerability Patching
por: Mastropaolo, Antonio, et al.
Publicado: (2024)
por: Mastropaolo, Antonio, et al.
Publicado: (2024)
Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
por: Yan, Yanfu, et al.
Publicado: (2025)
por: Yan, Yanfu, et al.
Publicado: (2025)
Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
Automated categorization of pre-trained models for software engineering: A case study with a Hugging Face dataset
por: Di Sipio, Claudio, et al.
Publicado: (2024)
por: Di Sipio, Claudio, et al.
Publicado: (2024)
Exploring the Role of Women in Hugging Face Organizations
por: Salinas, Maria Tubella, et al.
Publicado: (2025)
por: Salinas, Maria Tubella, et al.
Publicado: (2025)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
por: Taraghi, Mina, et al.
Publicado: (2024)
por: Taraghi, Mina, et al.
Publicado: (2024)
Rethinking Software Empirical Studies with Structural Causal Models
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2026)
"Let it be Chaos in the Plumbing!" Usage and Efficacy of Chaos Engineering in DevOps Pipelines
por: Fossati, Stefano, et al.
Publicado: (2025)
por: Fossati, Stefano, et al.
Publicado: (2025)
Testing Practices, Challenges, and Developer Perspectives in Open-Source IoT Platforms
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
Lessons Learned from Mining the Hugging Face Repository
por: Castaño, Joel, et al.
Publicado: (2024)
por: Castaño, Joel, et al.
Publicado: (2024)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot
por: Bifolco, Daniele, et al.
Publicado: (2025)
por: Bifolco, Daniele, et al.
Publicado: (2025)
Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
por: Giagnorio, Alessandro, et al.
Publicado: (2025)
por: Giagnorio, Alessandro, et al.
Publicado: (2025)
A Taxonomy of Self-Admitted Technical Debt in Deep Learning Systems
por: Pepe, Federica, et al.
Publicado: (2024)
por: Pepe, Federica, et al.
Publicado: (2024)
Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
por: Velasco, Alejandro, et al.
Publicado: (2024)
por: Velasco, Alejandro, et al.
Publicado: (2024)
Analyzing the Evolution and Maintenance of ML Models on Hugging Face
por: Castaño, Joel, et al.
Publicado: (2023)
por: Castaño, Joel, et al.
Publicado: (2023)
Ejemplares similares
-
"The Law Doesn't Work Like a Computer": Exploring Software Licensing Issues Faced by Legal Practitioners
por: Wintersgill, Nathan, et al.
Publicado: (2024) -
Developers' Perspectives on Software Licensing: Current Practices, Challenges, and Tools
por: Wintersgill, Nathan, et al.
Publicado: (2025) -
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
por: Stalnaker, Trevor, et al.
Publicado: (2024) -
BOMs Away! Inside the Minds of Stakeholders: A Comprehensive Study of Bills of Materials for Software Systems
por: Stalnaker, Trevor, et al.
Publicado: (2023) -
Prompting in Practice: Investigating Software Practitioners' Use of Generative AI Tools
por: Otten, Daniel, et al.
Publicado: (2025)