Detecting Malicious Source Code in PyPI Packages with LLMs: Does RAG Come in Handy?
Fuente:
arXiv
Saved in:
| Main Authors: | Ibiyo, Motunrayo, Louangdy, Thinakone, Nguyen, Phuong T., Di Sipio, Claudio, Di Ruscio, Davide |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Many Hands Make Light Work: An LLM-based Multi-Agent System for Detecting Malicious PyPI Packages
by: Zeshan, Muhammad Umar, et al.
Published: (2026)
by: Zeshan, Muhammad Umar, et al.
Published: (2026)
A Machine Learning-Based Approach For Detecting Malicious PyPI Packages
by: Samaana, Haya, et al.
Published: (2024)
by: Samaana, Haya, et al.
Published: (2024)
CHASE: LLM Agents for Dissecting Malicious PyPI Packages
by: Toda, Takaaki, et al.
Published: (2026)
by: Toda, Takaaki, et al.
Published: (2026)
On the use of Large Language Models in Model-Driven Engineering
by: Di Rocco, Juri, et al.
Published: (2024)
by: Di Rocco, Juri, et al.
Published: (2024)
Automated categorization of pre-trained models for software engineering: A case study with a Hugging Face dataset
by: Di Sipio, Claudio, et al.
Published: (2024)
by: Di Sipio, Claudio, et al.
Published: (2024)
Addressing Popularity Bias in Third-Party Library Recommendations Using LLMs
by: Di Sipio, Claudio, et al.
Published: (2025)
by: Di Sipio, Claudio, et al.
Published: (2025)
PyRadar: Towards Automatically Retrieving and Validating Source Code Repository Information for PyPI Packages
by: Gao, Kai, et al.
Published: (2024)
by: Gao, Kai, et al.
Published: (2024)
Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining Framework
by: Guo, Wenbo, et al.
Published: (2026)
by: Guo, Wenbo, et al.
Published: (2026)
Automatic Categorization of GitHub Actions with Transformers and Few-shot Learning
by: Nguyen, Phuong T., et al.
Published: (2024)
by: Nguyen, Phuong T., et al.
Published: (2024)
DySec: A Machine Learning-based Dynamic Analysis for Detecting Malicious Packages in PyPI Ecosystem
by: Mehedi, Sk Tanzir, et al.
Published: (2025)
by: Mehedi, Sk Tanzir, et al.
Published: (2025)
Detection of Technical Debt in Java Source Code
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
Investigating the Role of LLMs Hyperparameter Tuning and Prompt Engineering to Support Domain Modeling
by: Bulhakov, Vladyslav, et al.
Published: (2025)
by: Bulhakov, Vladyslav, et al.
Published: (2025)
Killing Two Birds with One Stone: Malicious Package Detection in NPM and PyPI using a Single Model of Malicious Behavior Sequence
by: Zhang, Junan, et al.
Published: (2023)
by: Zhang, Junan, et al.
Published: (2023)
How fair are we? From conceptualization to automated assessment of fairness definitions
by: d'Aloisio, Giordano, et al.
Published: (2024)
by: d'Aloisio, Giordano, et al.
Published: (2024)
When simplicity meets effectiveness: Detecting code comments coherence with word embeddings and LSTM
by: Igbomezie, Michael Dubem, et al.
Published: (2024)
by: Igbomezie, Michael Dubem, et al.
Published: (2024)
Simplicity by Obfuscation: Evaluating LLM-Driven Code Transformation with Semantic Elasticity
by: De Tomasi, Lorenzo, et al.
Published: (2025)
by: De Tomasi, Lorenzo, et al.
Published: (2025)
SourceBroken: A large-scale analysis on the (un)reliability of SourceRank in the PyPI ecosystem
by: Montaruli, Biagio, et al.
Published: (2025)
by: Montaruli, Biagio, et al.
Published: (2025)
Teamwork makes the dream work: LLMs-Based Agents for GitHub README.MD Summarization
by: Nguyen, Duc S. H., et al.
Published: (2025)
by: Nguyen, Duc S. H., et al.
Published: (2025)
Analyzing the Usage of Donation Platforms for PyPI Libraries
by: Tsakpinis, Alexandros, et al.
Published: (2025)
by: Tsakpinis, Alexandros, et al.
Published: (2025)
How Maintainable is Proficient Code? A Case Study of Three PyPI Libraries
by: Febriyanti, Indira, et al.
Published: (2024)
by: Febriyanti, Indira, et al.
Published: (2024)
Analyzing the Availability of E-Mail Addresses for PyPI Libraries
by: Tsakpinis, Alexandros, et al.
Published: (2026)
by: Tsakpinis, Alexandros, et al.
Published: (2026)
Prompt engineering and its implications on the energy consumption of Large Language Models
by: Rubei, Riccardo, et al.
Published: (2025)
by: Rubei, Riccardo, et al.
Published: (2025)
Analyzing the Accessibility of GitHub Repositories for PyPI and NPM Libraries
by: Tsakpinis, Alexandros, et al.
Published: (2024)
by: Tsakpinis, Alexandros, et al.
Published: (2024)
Forecasting the Maintained Score from the OpenSSF Scorecard: A Study of GitHub Repositories Linked to PyPI Packages
by: Tsakpinis, Alexandros, et al.
Published: (2026)
by: Tsakpinis, Alexandros, et al.
Published: (2026)
Exploring the SECURITY.md in the Dependency Chain: Preliminary Analysis of the PyPI Ecosystem
by: Termphaiboon, Chayanid, et al.
Published: (2025)
by: Termphaiboon, Chayanid, et al.
Published: (2025)
Less is More? An Empirical Study on Configuration Issues in Python PyPI Ecosystem
by: Peng, Yun, et al.
Published: (2023)
by: Peng, Yun, et al.
Published: (2023)
Bake Two Cakes with One Oven: RL for Defusing Popularity Bias and Cold-start in Third-Party Library Recommendations
by: Vuong, Minh Hoang, et al.
Published: (2025)
by: Vuong, Minh Hoang, et al.
Published: (2025)
Good things come in three: Generating SO Post Titles with Pre-Trained Models, Self Improvement and Post Ranking
by: Le, Duc Anh, et al.
Published: (2024)
by: Le, Duc Anh, et al.
Published: (2024)
Small Changes, Big Trouble: Demystifying and Parsing License Variants for Incompatibility Detection in the PyPI Ecosystem
by: Xu, Weiwei, et al.
Published: (2025)
by: Xu, Weiwei, et al.
Published: (2025)
MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI Ecosystem
by: Gao, Xingan, et al.
Published: (2025)
by: Gao, Xingan, et al.
Published: (2025)
One Detector Fits All: Robust and Adaptive Detection of Malicious Packages from PyPI to Enterprises
by: Montaruli, Biagio, et al.
Published: (2025)
by: Montaruli, Biagio, et al.
Published: (2025)
Investigating Notable Metadata Practices in PyPI Libraries: An Empirical Study about Repository and Donation Platform URLs
by: Tsakpinis, Alexandros, et al.
Published: (2026)
by: Tsakpinis, Alexandros, et al.
Published: (2026)
Modeling Dependency-Propagated Ecosystem Impact of Changes in Maintenance Activities: Evaluating Support Strategies in the PyPI Network
by: Tsakpinis, Alexandros, et al.
Published: (2026)
by: Tsakpinis, Alexandros, et al.
Published: (2026)
Automation in Model-Driven Engineering: A look back, and ahead
by: Burgueño, Lola, et al.
Published: (2024)
by: Burgueño, Lola, et al.
Published: (2024)
An Analysis of Malicious Packages in Open-Source Software in the Wild
by: Zhou, Xiaoyan, et al.
Published: (2024)
by: Zhou, Xiaoyan, et al.
Published: (2024)
PlayMyData: a curated dataset of multi-platform video games
by: D'Angelo, Andrea, et al.
Published: (2024)
by: D'Angelo, Andrea, et al.
Published: (2024)
Advanced discovery mechanisms in model repositories
by: Arsene Indamutsa, et al.
Published: (2024)
by: Arsene Indamutsa, et al.
Published: (2024)
On the Need for Configurable Travel Recommender Systems: A Systematic Mapping Study
by: Pereira, Rickson Simioni, et al.
Published: (2024)
by: Pereira, Rickson Simioni, et al.
Published: (2024)
Towards Synthetic Trace Generation of Modeling Operations using In-Context Learning Approach
by: Muttillo, Vittoriano, et al.
Published: (2024)
by: Muttillo, Vittoriano, et al.
Published: (2024)
Larger Is Not Always Better: Leveraging Structured Code Diffs for Comment Inconsistency Detection
by: Nguyen, Phong, et al.
Published: (2025)
by: Nguyen, Phong, et al.
Published: (2025)
Similar Items
-
Many Hands Make Light Work: An LLM-based Multi-Agent System for Detecting Malicious PyPI Packages
by: Zeshan, Muhammad Umar, et al.
Published: (2026) -
A Machine Learning-Based Approach For Detecting Malicious PyPI Packages
by: Samaana, Haya, et al.
Published: (2024) -
CHASE: LLM Agents for Dissecting Malicious PyPI Packages
by: Toda, Takaaki, et al.
Published: (2026) -
On the use of Large Language Models in Model-Driven Engineering
by: Di Rocco, Juri, et al.
Published: (2024) -
Automated categorization of pre-trained models for software engineering: A case study with a Hugging Face dataset
by: Di Sipio, Claudio, et al.
Published: (2024)