Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zihan, Zhang, Rui, Liu, Yu, Liu, Chi, Zhao, Qingchuan, Li, Hongwei, Xu, Guowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MPMA: Preference Manipulation Attack Against Model Context Protocol
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
Combinational Backdoor Attack against Customized Text-to-Image Models
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
von: Li, Jiate, et al.
Veröffentlicht: (2026)
von: Li, Jiate, et al.
Veröffentlicht: (2026)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
Black-Box Guardrail Reverse-engineering Attack
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
von: Yao, Hongwei, et al.
Veröffentlicht: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
von: Wang, Haowei, et al.
Veröffentlicht: (2025)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
von: Wang, Xinkai, et al.
Veröffentlicht: (2025)
OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud
von: Yuan, Shuai, et al.
Veröffentlicht: (2024)
von: Yuan, Shuai, et al.
Veröffentlicht: (2024)
A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
von: Liu, Zixuan, et al.
Veröffentlicht: (2025)
Assessing Risk of Stealing Proprietary Models for Medical Imaging Tasks
von: Raj, Ankita, et al.
Veröffentlicht: (2025)
von: Raj, Ankita, et al.
Veröffentlicht: (2025)
Efficient Data-Free Model Stealing with Label Diversity
von: Liu, Yiyong, et al.
Veröffentlicht: (2024)
von: Liu, Yiyong, et al.
Veröffentlicht: (2024)
Backdoor Attacks against Image-to-Image Networks
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
von: Zhao, Shiqian, et al.
Veröffentlicht: (2025)
von: Zhao, Shiqian, et al.
Veröffentlicht: (2025)
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills
von: Liu, Yiyong, et al.
Veröffentlicht: (2026)
von: Liu, Yiyong, et al.
Veröffentlicht: (2026)
GRID: Protecting Training Graph from Link Stealing Attacks on GNN Models
von: Lou, Jiadong, et al.
Veröffentlicht: (2025)
von: Lou, Jiadong, et al.
Veröffentlicht: (2025)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
von: Liu, Yi, et al.
Veröffentlicht: (2026)
von: Liu, Yi, et al.
Veröffentlicht: (2026)
InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks
von: Zheng, Xinyao, et al.
Veröffentlicht: (2024)
von: Zheng, Xinyao, et al.
Veröffentlicht: (2024)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
von: Domico, Kyle, et al.
Veröffentlicht: (2025)
von: Domico, Kyle, et al.
Veröffentlicht: (2025)
Model Stealing Attack against Graph Classification with Authenticity, Uncertainty and Diversity
von: Zhu, Zhihao, et al.
Veröffentlicht: (2023)
von: Zhu, Zhihao, et al.
Veröffentlicht: (2023)
BESA: Boosting Encoder Stealing Attack with Perturbation Recovery
von: Ren, Xuhao, et al.
Veröffentlicht: (2025)
von: Ren, Xuhao, et al.
Veröffentlicht: (2025)
Stealing Trust: Unraveling Blind Message Attacks in Web3 Authentication
von: Yan, Kailun, et al.
Veröffentlicht: (2024)
von: Yan, Kailun, et al.
Veröffentlicht: (2024)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
von: Duan, Zenghao, et al.
Veröffentlicht: (2026)
Prompt Stealing Attacks Against Large Language Models
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
von: Sha, Zeyang, et al.
Veröffentlicht: (2024)
When Skills Lie: Hidden-Comment Injection in LLM Agents
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills
von: Li, Ying, et al.
Veröffentlicht: (2026)
von: Li, Ying, et al.
Veröffentlicht: (2026)
A Model Stealing Attack Against Multi-Exit Networks
von: Pan, Li, et al.
Veröffentlicht: (2023)
von: Pan, Li, et al.
Veröffentlicht: (2023)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
von: Wang, Weijie, et al.
Veröffentlicht: (2026)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MPMA: Preference Manipulation Attack Against Model Context Protocol
von: Wang, Zihan, et al.
Veröffentlicht: (2025) -
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025) -
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025) -
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026) -
Combinational Backdoor Attack against Customized Text-to-Image Models
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)