Evolving Deception: When Agents Evolve, Deception Wins
Fuente:
arXiv
Salvato in:
| Autori principali: | Ying, Zonghao, Dai, Haowen, Zhang, Tianyuan, Xiao, Yisong, Zou, Quanchen, Liu, Aishan, Yang, Jian, Yang, Yaodong, Liu, Xianglong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
di: Zou, Quanchen, et al.
Pubblicazione: (2026)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
di: Zou, Quanchen, et al.
Pubblicazione: (2025)
di: Zou, Quanchen, et al.
Pubblicazione: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
di: Liu, Zhe, et al.
Pubblicazione: (2026)
di: Liu, Zhe, et al.
Pubblicazione: (2026)
WiP: Deception-in-Depth Using Multiple Layers of Deception
di: Landsborough, Jason, et al.
Pubblicazione: (2024)
di: Landsborough, Jason, et al.
Pubblicazione: (2024)
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
di: Xu, Wenzhuo, et al.
Pubblicazione: (2026)
Baiting AI: Deceptive Adversary Against AI-Protected Industrial Infrastructures
di: Pasikhani, Aryan, et al.
Pubblicazione: (2026)
di: Pasikhani, Aryan, et al.
Pubblicazione: (2026)
Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars
di: Tehrani, Masoud Jamshidiyan, et al.
Pubblicazione: (2026)
di: Tehrani, Masoud Jamshidiyan, et al.
Pubblicazione: (2026)
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models
di: Yang, Ying, et al.
Pubblicazione: (2025)
di: Yang, Ying, et al.
Pubblicazione: (2025)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
di: Xu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Xu, Wenzhuo, et al.
Pubblicazione: (2025)
Illusion Worlds: Deceptive UI Attacks in Social VR
di: Lee, Junhee, et al.
Pubblicazione: (2025)
di: Lee, Junhee, et al.
Pubblicazione: (2025)
Resource-aware Cyber Deception for Microservice-based Applications
di: Zambianco, Marco, et al.
Pubblicazione: (2023)
di: Zambianco, Marco, et al.
Pubblicazione: (2023)
Koney: A Cyber Deception Orchestration Framework for Kubernetes
di: Kahlhofer, Mario, et al.
Pubblicazione: (2025)
di: Kahlhofer, Mario, et al.
Pubblicazione: (2025)
AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
di: Rassul, Yassin H., et al.
Pubblicazione: (2026)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2024)
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2024)
Access Over Deception: Fighting Deceptive Patterns through Accessibility
di: Pellkvist, Tobias, et al.
Pubblicazione: (2026)
di: Pellkvist, Tobias, et al.
Pubblicazione: (2026)
PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
di: Zhang, Jianqing, et al.
Pubblicazione: (2025)
Physical Layer Deception in OFDM Systems
di: Chen, Wenwen, et al.
Pubblicazione: (2024)
di: Chen, Wenwen, et al.
Pubblicazione: (2024)
Coordinated Multi-Domain Deception: A Stackelberg Game Approach
di: Sayed, Md Abu, et al.
Pubblicazione: (2026)
di: Sayed, Md Abu, et al.
Pubblicazione: (2026)
A Survey of Network Requirements for Enabling Effective Cyber Deception
di: Sayed, Md Abu, et al.
Pubblicazione: (2023)
di: Sayed, Md Abu, et al.
Pubblicazione: (2023)
An Explainable XGBoost-based Approach on Assessing Detection of Deception and Disinformation
di: Mbaziira, Alex V, et al.
Pubblicazione: (2024)
di: Mbaziira, Alex V, et al.
Pubblicazione: (2024)
ranDecepter: Real-time Identification and Deterrence of Ransomware Attacks
di: Sajid, Md Sajidul Islam, et al.
Pubblicazione: (2025)
di: Sajid, Md Sajidul Islam, et al.
Pubblicazione: (2025)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
di: Li, Jindong, et al.
Pubblicazione: (2026)
di: Li, Jindong, et al.
Pubblicazione: (2026)
50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security Implications
di: Shi, Zewei, et al.
Pubblicazione: (2025)
di: Shi, Zewei, et al.
Pubblicazione: (2025)
Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
When AI Defeats Password Deception! A Deep Learning Framework to Distinguish Passwords and Honeywords
di: Dani, Jimmy, et al.
Pubblicazione: (2024)
di: Dani, Jimmy, et al.
Pubblicazione: (2024)
Perry: A High-level Framework for Accelerating Cyber Deception Experimentation
di: Singer, Brian, et al.
Pubblicazione: (2025)
di: Singer, Brian, et al.
Pubblicazione: (2025)
VENENA: A Deceptive Visual Encryption Framework for Wireless Semantic Secrecy
di: Han, Bin, et al.
Pubblicazione: (2025)
di: Han, Bin, et al.
Pubblicazione: (2025)
The Invisible Game on the Internet: A Case Study of Decoding Deceptive Patterns
di: Shi, Zewei, et al.
Pubblicazione: (2024)
di: Shi, Zewei, et al.
Pubblicazione: (2024)
A Descriptive Model for Modelling Attacker Decision-Making in Cyber-Deception
di: Turner, B. R., et al.
Pubblicazione: (2025)
di: Turner, B. R., et al.
Pubblicazione: (2025)
Documenti analoghi
-
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026) -
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
di: Ying, Zonghao, et al.
Pubblicazione: (2025) -
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
di: Ying, Zonghao, et al.
Pubblicazione: (2025) -
Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models
di: Zou, Quanchen, et al.
Pubblicazione: (2026) -
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
di: Ying, Zonghao, et al.
Pubblicazione: (2024)