SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Jianshuo, Guo, Sheng, Wang, Hao, Chen, Xun, Liu, Zhuotao, Zhang, Tianwei, Xu, Ke, Huang, Minlie, Qiu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
by: Ou, Haoran, et al.
Published: (2025)
by: Ou, Haoran, et al.
Published: (2025)
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Can Large Language Models Automate the Refinement of Cellular Network Specifications?
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)
by: Li, Boheng, et al.
Published: (2025)
An Engorgio Prompt Makes Large Language Model Babble on
by: Dong, Jianshuo, et al.
Published: (2024)
by: Dong, Jianshuo, et al.
Published: (2024)
SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
by: Duan, Zenghao, et al.
Published: (2026)
by: Duan, Zenghao, et al.
Published: (2026)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
by: Jotautaitė, Monika, et al.
Published: (2026)
by: Jotautaitė, Monika, et al.
Published: (2026)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
LeakDojo: Decoding the Leakage Threats of RAG Systems
by: Zhang, Maosen, et al.
Published: (2026)
by: Zhang, Maosen, et al.
Published: (2026)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
by: Wei, Zhang, et al.
Published: (2025)
by: Wei, Zhang, et al.
Published: (2025)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
martFL: Enabling Utility-Driven Data Marketplace with a Robust and Verifiable Federated Learning Architecture
by: Li, Qi, et al.
Published: (2023)
by: Li, Qi, et al.
Published: (2023)
ObfusBFA: A Holistic Approach to Safeguarding DNNs from Different Types of Bit-Flip Attacks
by: Yan, Xiaobei, et al.
Published: (2025)
by: Yan, Xiaobei, et al.
Published: (2025)
A Systematic Study of LLM-Based Architectures for Automated Patching
by: Xu, Qingxiao, et al.
Published: (2026)
by: Xu, Qingxiao, et al.
Published: (2026)
Red Teaming Methodology for Design Obfuscation
by: Liu, Yuntao, et al.
Published: (2025)
by: Liu, Yuntao, et al.
Published: (2025)
Towards Understanding the Cognitive Habits of Large Reasoning Models
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
by: Fumero, Stefano, et al.
Published: (2025)
by: Fumero, Stefano, et al.
Published: (2025)
Tricking LLM-Based NPCs into Spilling Secrets
by: Shiomi, Kyohei, et al.
Published: (2025)
by: Shiomi, Kyohei, et al.
Published: (2025)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
RulePilot: An LLM-Powered Agent for Security Rule Generation
by: Wang, Hongtai, et al.
Published: (2025)
by: Wang, Hongtai, et al.
Published: (2025)
A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
FuzzingBrain V2: A Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction
by: Sheng, Ze, et al.
Published: (2026)
by: Sheng, Ze, et al.
Published: (2026)
RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
by: Huang, Hanbo, et al.
Published: (2025)
by: Huang, Hanbo, et al.
Published: (2025)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026)
by: Ou, Haoran, et al.
Published: (2026)
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
by: Xu, Huiyu, et al.
Published: (2024)
by: Xu, Huiyu, et al.
Published: (2024)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
All You Need Is A Fuzzing Brain: An LLM-Powered System for Automated Vulnerability Detection and Patching
by: Sheng, Ze, et al.
Published: (2025)
by: Sheng, Ze, et al.
Published: (2025)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
by: Shen, Xiangmin, et al.
Published: (2024)
by: Shen, Xiangmin, et al.
Published: (2024)
LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems
by: Atta, Hammad, et al.
Published: (2026)
by: Atta, Hammad, et al.
Published: (2026)
Pencil: Private and Extensible Collaborative Learning without the Non-Colluding Assumption
by: Liu, Xuanqi, et al.
Published: (2024)
by: Liu, Xuanqi, et al.
Published: (2024)
Similar Items
-
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
by: Ou, Haoran, et al.
Published: (2025) -
Red-Teaming LLM Multi-Agent Systems via Communication Attacks
by: He, Pengfei, et al.
Published: (2025) -
Can Large Language Models Automate the Refinement of Cellular Network Specifications?
by: Dong, Jianshuo, et al.
Published: (2025) -
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026) -
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
by: Li, Boheng, et al.
Published: (2025)