Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yi, Wang, Weizhe, Feng, Ruitao, Zhang, Yao, Xu, Guangquan, Deng, Gelei, Li, Yuekang, Zhang, Leo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026)
by: Qu, Yubin, et al.
Published: (2026)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
An Empirical Study on the Security Vulnerabilities of GPTs
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
LLM-Enabled Open-Source Systems in the Wild: An Empirical Study of Vulnerabilities in GitHub Security Advisories
by: Shifat, Fariha Tanjim, et al.
Published: (2026)
by: Shifat, Fariha Tanjim, et al.
Published: (2026)
FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
by: Zhou, Zhiping, et al.
Published: (2025)
by: Zhou, Zhiping, et al.
Published: (2025)
SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysis
by: Lingxiang, Wang, et al.
Published: (2025)
by: Lingxiang, Wang, et al.
Published: (2025)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Unlocking User-oriented Pages: Intention-driven Black-box Scanner for Real-world Web Applications
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
Characterizing Trust Boundary Vulnerabilities in TEE Containers: An Empirical Study
by: Liu, Weijie, et al.
Published: (2025)
by: Liu, Weijie, et al.
Published: (2025)
Argus: Reorchestrating Static Analysis via a Multi-Agent Ensemble for Full-Chain Security Vulnerability Detection
by: Liang, Zi, et al.
Published: (2026)
by: Liang, Zi, et al.
Published: (2026)
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
by: Guo, Zihan, et al.
Published: (2026)
by: Guo, Zihan, et al.
Published: (2026)
A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and Practitioners
by: Shi, Jingyi, et al.
Published: (2025)
by: Shi, Jingyi, et al.
Published: (2025)
IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
by: Li, Ziyang, et al.
Published: (2024)
by: Li, Ziyang, et al.
Published: (2024)
Understanding the Effectiveness of Large Language Models in Detecting Security Vulnerabilities
by: Khare, Avishree, et al.
Published: (2023)
by: Khare, Avishree, et al.
Published: (2023)
QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
by: Wang, Claire, et al.
Published: (2025)
by: Wang, Claire, et al.
Published: (2025)
Revisiting Vulnerability Patch Identification on Data in the Wild
by: Irsan, Ivana Clairine, et al.
Published: (2026)
by: Irsan, Ivana Clairine, et al.
Published: (2026)
Unity is Strength: Enhancing Precision in Reentrancy Vulnerability Detection of Smart Contract Analysis Tools
by: Wang, Zexu, et al.
Published: (2024)
by: Wang, Zexu, et al.
Published: (2024)
An Empirical Study of Vulnerability Handling Times in CPython
by: Ruohonen, Jukka
Published: (2024)
by: Ruohonen, Jukka
Published: (2024)
Generating Proof-of-Vulnerability Tests to Help Enhance the Security of Complex Software
by: Kanchi, Shravya, et al.
Published: (2026)
by: Kanchi, Shravya, et al.
Published: (2026)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
by: Liu, Shuhan, et al.
Published: (2025)
by: Liu, Shuhan, et al.
Published: (2025)
One Signature, Multiple Payments: Demystifying and Detecting Signature Replay Vulnerabilities in Smart Contracts
by: Wang, Zexu, et al.
Published: (2025)
by: Wang, Zexu, et al.
Published: (2025)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
by: Xu, Xiangzhe, et al.
Published: (2024)
by: Xu, Xiangzhe, et al.
Published: (2024)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Security Incentivization: An Empirical Study of how Micropayments Impact Code Security
by: Rass, Stefan, et al.
Published: (2026)
by: Rass, Stefan, et al.
Published: (2026)
Conflicting Scores, Confusing Signals: An Empirical Study of Vulnerability Scoring Systems
by: Koscinski, Viktoria, et al.
Published: (2025)
by: Koscinski, Viktoria, et al.
Published: (2025)
An Empirical Study on Oculus Virtual Reality Applications: Security and Privacy Perspectives
by: Guo, Hanyang, et al.
Published: (2024)
by: Guo, Hanyang, et al.
Published: (2024)
Fixing Security Vulnerabilities with AI in OSS-Fuzz
by: Zhang, Yuntong, et al.
Published: (2024)
by: Zhang, Yuntong, et al.
Published: (2024)
On the Security Vulnerabilities of Text-to-SQL Models
by: Peng, Xutan, et al.
Published: (2022)
by: Peng, Xutan, et al.
Published: (2022)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
Exploring Security Practices in Infrastructure as Code: An Empirical Study
by: Verdet, Alexandre, et al.
Published: (2023)
by: Verdet, Alexandre, et al.
Published: (2023)
Efficient Detection of Toxic Prompts in Large Language Models
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
An Empirical Study of Bitwise Operators Intuitiveness through Performance Metrics
by: Joshi, Shubham
Published: (2025)
by: Joshi, Shubham
Published: (2025)
A Large-scale Empirical Study on the Generalizability of Disclosed Java Library Vulnerability Exploits
by: Chen, Zirui, et al.
Published: (2026)
by: Chen, Zirui, et al.
Published: (2026)
WildCode: An Empirical Analysis of Code Generated by ChatGPT
by: Khanmohammadi, Kobra, et al.
Published: (2025)
by: Khanmohammadi, Kobra, et al.
Published: (2025)
The Secrets Must Not Flow: Scaling Security Verification to Large Codebases (extended version)
by: Arquint, Linard, et al.
Published: (2025)
by: Arquint, Linard, et al.
Published: (2025)
Security study based on the Chatgptplugin system: ldentifying Security Vulnerabilities
by: Ren, Ruomai
Published: (2025)
by: Ren, Ruomai
Published: (2025)
Similar Items
-
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
by: Qu, Yubin, et al.
Published: (2026) -
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026) -
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025) -
An Empirical Study on the Security Vulnerabilities of GPTs
by: Wu, Tong, et al.
Published: (2025) -
LLM-Enabled Open-Source Systems in the Wild: An Empirical Study of Vulnerabilities in GitHub Security Advisories
by: Shifat, Fariha Tanjim, et al.
Published: (2026)